Journal · August 25, 2026

Planning-model evaluation: a practical working assessment

This is a concise record of a small planning experiment and the operating approach it informed. It is not a general benchmark.

Method

A small qualitative planning experiment.

I used OpenCode as the harness and assessed the quality of eight submitted plans. I also attempted local Ollama runs with Qwen Coder and DeepSeek Coder models above 10B parameters on a 2.4 GHz quad-core Intel Core i5 laptop with 8 GB LPDDR3 memory.

The comparison also included OpenAI API tests using GPT-5.6 Luna and GPT-5.6 Terra, free OpenCode model access, and Amazon Bedrock. This is a record of a practical experiment, not a provider-wide or statistically valid benchmark.

Finding

Luna ranked first and Terra second for the submitted plans.

In this eight-plan assessment, GPT-5.6 Luna ranked first and GPT-5.6 Terra ranked second for plan quality. That result is a personal observation from this limited evaluation, not a claim that either model will rank the same way for another workflow or team.

I did not record cost, token-usage, or elapsed-time artifacts. The evaluation therefore cannot support a cost, value-for-money, or speed conclusion.

Working approach

Use available capacity thoughtfully and keep alternatives practical.

My current working assessment is to make the most of existing Codex credits, prefer GPT-5.6 Terra and Luna for planning, and keep Amazon Bedrock plus Anthropic Sonnet or DeepSeek as practical alternatives when they fit the work. This is a personal operating approach, not advice about provider limits, pricing, or availability.

Codex can feel more capable in practice because it combines a coding-optimized model with an agent harness, workspace context, tools, permissions, and workflow integration. That combined experience is different from model capability alone; it does not show that an underlying Codex model is categorically superior to API or Bedrock models.