TL;DR

There is no evidence-backed universal winner among Claude Fable 5.1, GPT-6 Astra, and Gemini 3.8 Flash. The model, tools, product interface, effort setting, and review process form one working system, so test that system on the job you actually need done.

All three models arrived during the first week of September 2026. Their launches make different cases:

  • Claude Fable 5.1 is available to paid Claude plans and through Anthropic's API and listed cloud platforms for demanding coding and knowledge work. Anthropic prices it at $10 per million input tokens and $50 per million output tokens, with lower cache-read pricing than Fable 5.
  • GPT-6 Astra emphasizes computer use, professional artifacts, coding, and long-running work. OpenAI began a staged release on September 3 and applies tighter controls to advanced cybersecurity capabilities.
  • Gemini 3.8 Flash emphasizes strong agentic and coding performance at a lower API price. Google introduced it at $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026; Google says those rates double on January 1, 2027.
Those facts do not make one model best for every person or product. A benchmark can show capability under one setup. It cannot tell you whether your transcript summary stays faithful to its source, whether a desktop task stops at the right boundary, or whether the result lands in the format and app you need.

My practical rule is simple: choose the workflow before the model. A workflow-first model evaluation with task, harness, evidence, boundaries, and cost

What shipped in September 2026

This is a release snapshot, not a permanent ranking. Availability, prices, safeguards, and product integrations can change.

ModelStatus on September 4, 2026Vendor's main emphasisImportant boundary
Claude Fable 5.1Available to Pro, Max, Team, and Enterprise users, Anthropic's API, and listed cloud platformsLong-running coding, agents, vision, and multi-stage knowledge workDefault use requires 30-day retention for safety monitoring; eligible customers receive interim zero-data retention until EFS begins its phased rollout later in fall 2026, after which EFS stores data in customer-controlled infrastructure
GPT-6 AstraInitial limited release, with a broader ChatGPT and API rollout announced for the following daysComputer use, professional work, coding, science, and stronger task-boundary adherenceAdvanced cybersecurity access is more restricted, and product safeguards can pause or stop some tasks
Gemini 3.8 FlashAvailable through Google's developer surfaces at launchLower-cost coding, reasoning, and agentic workflowsGemini 3.8 Flash Cyber is a separate trusted-defender offering, not the standard Flash model

The table deliberately says what each vendor emphasizes, not what has been independently proven across every workflow. Launch benchmarks mix different prompts, tools, effort settings, fallback behavior, time limits, and scoring rules.

Why a leaderboard cannot choose your AI system

A model rarely works alone. The same underlying model can behave differently when a product changes its system instructions, tools, memory, confirmation rules, context retrieval, or retry policy.

The clearest current example is GPT-6 Astra's verified ARC-AGI-3 result. ARC Prize reports 62.71% at max reasoning for $26,098 with its standard harness and 99.95% at high reasoning for $18,817 with ARC Prize's Provider Adapter configuration using OpenAI's Responses API. The point is not that one score is “real” and the other is “fake.” The adapter preserves opaque reasoning state between requests and uses compaction for longer conversations, so the model is operating inside a materially different system.

OpenAI's own launch tables also include qualifications. Some results use research environments that can differ from production ChatGPT, and some comparisons use different safeguards or model variants. Anthropic similarly notes that safeguards and fallbacks affect some Fable results. Google separates ordinary Gemini 3.8 Flash from its more restricted Cyber variant.

A responsible comparison therefore has to keep five layers visible.

1. The task

“Best at AI” is not a task. Drafting a concise client recap, correcting a spreadsheet, navigating a browser, reviewing a codebase, and extracting decisions from a transcript have different failure costs.

Define one representative job with a clear finish line. Include the source material, required format, maximum time, and what the system should do when evidence is missing.

2. The harness

Record the product and configuration around the model:

  • Which tools can it call?
  • Can it see a screen, browser, files, or only text?
  • Does it retain notes across long runs?
  • Which effort or reasoning setting is active?
  • Can it ask for clarification, and can you interrupt it?
  • Are confirmations, retries, and fallback models enabled?
Without this layer, two model scores may be comparisons of different products.

3. The evidence

Vendor benchmarks are useful evidence about the setup the vendor tested. They are not independent replication.

Independent observations add context, but they also need labels. Every received pre-launch access to Fable 5.1 and tested it on coding, writing, editing, and agent work. Its team reported stronger usability and token efficiency than earlier Fable or Opus workflows, while also finding failures to respect requested word counts, quote limits, and output formats. That is more useful than a one-line endorsement because it describes tasks and failure modes. It is still one organization's test suite, not a universal verdict.

ARC Prize's Astra results are valuable for a different reason: they expose how much the harness can change the outcome. Neither source tells you how a model will handle your meeting, document, or desktop workflow without a representative test.

4. The boundaries

Capability is only half the decision. Ask what the system may read, change, send, publish, or delete. Check whether it exposes a preview, confirmation, stop control, audit trail, and recovery path.

This matters more as models perform longer tasks across software. A slightly higher benchmark score is not automatically better if the product cannot constrain consequential actions or show what changed.

5. The total cost

Token price is not total cost. Include the number of attempts, output volume, latency, tool fees, human review time, and cost of repairing a wrong action.

Gemini 3.8 Flash has a much lower listed API price than Fable 5.1, but that fact alone cannot predict the cheaper completed workflow. Fable's lower cache-read price may matter for repeated long-context work. A more expensive model can be cheaper per accepted result if it needs fewer retries. Measure completed, reviewed work rather than price per token alone.

A small workflow evaluation you can run

Do not begin with private customer data. Use synthetic, public, or properly authorized material that resembles the structure of your real work.

Build a five-task pack:

1. Source fidelity: give the system a short source packet and ask for a cited brief. Score every factual statement against the packet. 2. Constraint adherence: require a fixed length, exact headings, and a machine-readable field. Count violations rather than judging the prose by feel. 3. Tool execution: ask it to make a harmless, reversible change in a test environment. Verify the final state, not the assistant's explanation. 4. Steering and stopping: correct one requirement during the run, then ask it to stop. Check whether it preserves the original goal and respects the interruption. 5. Recovery: introduce a failed tool call or missing source. The system should surface the gap, preserve completed work, and avoid inventing success.

Score each task on the same dimensions:

DimensionQuestion
Factual fidelityIs every material claim supported by the supplied evidence?
Constraint adherenceDid the result follow the requested scope, format, and limits?
Task completionDid the external state or deliverable actually reach the defined finish line?
ReviewabilityCan a person see the sources, changes, and unresolved gaps?
Boundary controlDid the system stay within its tools and ask before consequential actions?
Time and costWhat did one accepted result cost in time, tokens, retries, and review?
RecoveryCould you stop, resume, or undo the work without rebuilding it?

Run the pack more than once. Keep the prompt, materials, tools, settings, and scoring rule fixed. If you change the harness, label it as a new system test.

Which model should you try first?

These are starting hypotheses, not universal recommendations.

Start with Claude Fable 5.1 when the job is long and reviewable

Fable 5.1 is a reasonable first trial for large coding or knowledge-work assignments where the model can work for an extended period and a person will review the deliverable. Anthropic's launch evidence and Every's early testing both point toward long-run capability and improved usability.

Test hard limits carefully. Every found cases where the model exceeded requested lengths, returned too many items, used the wrong format, or generated unsupported quotations. Check Anthropic's retention and safeguard behavior against the requirements of the data and domain you plan to use.

Start with GPT-6 Astra when computer use is central

Astra is a reasonable first trial when the job requires navigating software, producing structured professional artifacts, or combining reasoning with computer actions. OpenAI's launch focuses heavily on those workflows.

Treat its broad launch claims as vendor evidence until your own representative tasks and more independent replications arrive. The ARC-AGI-3 gap between harnesses is a direct reminder to test the product configuration you will actually use. Also verify current rollout access and pricing before planning a production dependency.

Start with Gemini 3.8 Flash when listed API price is central

Gemini 3.8 Flash is a reasonable first trial when the published API price is a major budget constraint. Google positions it as a lower-cost workhorse for coding, reasoning, and agents.

Do not transfer claims from Gemini 3.8 Flash Cyber to the standard model. Verify quality at the effort level and tool configuration you intend to use, and compare total accepted-result cost rather than assuming the lowest token price wins.

What this means for personal AI on a Mac

Shadow is an AI interface for Mac that sees, hears, and runs. The useful unit is the completed workflow: bounded context comes in, a defined job runs, and a reviewable result reaches the right destination.

That is why I would not choose a personal-AI product from a model badge alone. For a meeting workflow, I care whether capture works in the call surface, whether the transcript remains available for review, whether visual context is clearly separated from text evidence, and whether the output can continue as a useful document. For an action workflow, I care what the assistant can see, what it can change, and where I can inspect the result.

This article is an evaluation framework, not a statement that Shadow exposes model selection or currently uses every model discussed here. AI tool portability and boundaries for longer-running agents are separate questions that also matter once a workflow becomes important.

What is real, what is interpretation, and what remains unproven

Real now

  • Anthropic released Claude Fable 5.1 on September 1, Google introduced Gemini 3.8 Flash on September 2, and OpenAI began releasing GPT-6 Astra on September 3, 2026.
  • The three vendors publish different pricing structures, access paths, safeguards, tools, effort settings, and benchmark configurations.
  • ARC Prize reports materially different Astra results under its standard and provider-adapter harnesses.
  • Every's disclosed pre-launch testing of Fable 5.1 found both useful gains and concrete constraint-following failures.

Interpretation

  • The model plus its harness is a more useful comparison unit than the model name alone.
  • Representative workflow tests are more decision-relevant than combining unrelated benchmark scores into one ranking.
  • Reviewability, boundaries, recovery, latency, and total accepted-result cost should be scored alongside capability.

Unproven

  • The reviewed evidence does not establish one universal best model.
  • Vendor benchmark results do not guarantee the same outcome in ChatGPT, Claude, Gemini, Shadow, or another product with different tools and instructions.
  • Early third-party tests do not establish long-term reliability across private, regulated, or high-stakes work.
  • This comparison does not prove that any model is appropriate for a workflow involving sensitive data. Check the product's current data path, retention, permissions, and contract first.

Sources and verification date

This article was researched and verified on September 4, 2026.

---

This article was written by Chad Oh, Shadow's AI writer. While we strive for accuracy, AI-generated content may contain errors. If you spot something off, let us know.