All insights
Benchmark · 7 min read

A benchmark tie is not a model-selection decision

A worked assurance reading of the Fable 5.1 and GPT-6 Astra comparison: verify the version, separate independent evidence from vendor claims, and record the workload conditions that make a result meaningful.

By AI Assurance

A benchmark tie is not a procurement decision

Artificial Analysis currently places Claude Fable 5.1 and GPT-6 Astra level at 53 on its Intelligence Index v4.3. That is useful context, but it does not mean the models are interchangeable.

The index combines different tasks. A tie can hide meaningful differences in tool use, analysis, speed, cost and failure modes. Treat the result as a question to investigate, not an answer to buy.

Two glass evidence paths converge on a cobalt decision node
A model-selection record should connect published evidence to the decision and the workload that decision serves.

Verify the model and version first

Model names are not stable identifiers. Record the exact provider, model label, release date, access tier and API endpoint used in the test.

The provider pages identify Fable 5.1 as a 1 September 2026 release and Astra as a 3 September 2026 release.

A copied model name without its version is not reproducible evidence. See the primary release context from Anthropic and OpenAI.

Evidence hierarchy: independent, provider-run, practitioner

Start with an independent comparison where the methodology and task definitions are visible.

The Artificial Analysis comparison is the useful baseline for this worked example.

Provider-run tables remain valuable, but they answer a different question. They show what a provider measured under its chosen prompt, harness, safeguards and configuration.

Practitioner reports can reveal operational detail that formal tables omit. Use them as context, then verify the setup before treating a result as portable.

What the composite evaluations actually measure

The v4.3 index reports the same composite score for both models. Its component results are not the same: Astra leads on AutomationBench and Terminal-Bench, while Fable leads on GDPval, SciCode and Humanity's Last Exam.

53 / 53

Intelligence Index v4.3

Artificial Analysis

68% / 59%

AutomationBench-AA

Astra / Fable

59% / 52%

Terminal-Bench v4.0

Astra / Fable

This pattern suggests a practical distinction, not a universal ranking. Astra appears stronger on environment and tool-completion tasks. Fable appears stronger on several analytical and scientific tasks.

That is an inference from the published component pattern, not a claim that either model is categorically better.

The v4.3 methodology note explains the suite change and its limits.

Why benchmark rankings change

Version changes can move a ranking without either model changing. v4.3 added Terminal-Bench v4.0 and AutomationBench-AA, so its score should not be compared casually with an earlier index version.

Every record should therefore include the benchmark version, access date, task mix, scoring rule and any guardrail or completion rule. “The model scored 53” is incomplete.

Harness problem in coding-agent comparisons

Agent benchmarks measure a system: model, harness, tools, context policy, retry logic, permissions and stopping rule. A coding-agent tie may say as much about the harness as the model.

Artificial Analysis reports a tie in its Coding Agent Index when Astra runs in Codex and Fable runs in Claude Code. That is a useful operational comparison, but it is not a clean model-only experiment.

For an assurance file, record the harness, reasoning effort, tool permissions, context window, retries, time limit and whether the agent can inspect or modify its environment.

Cost is workload-dependent

The comparison reports similar list prices, but different measured economics. Astra's reported index cost per task is lower, while Fable's cache price is lower and its runs use more output tokens.

$3.26 / $7.63

Measured cost per index task

Astra / Fable

27k / 78k

Output tokens per task

Astra / Fable

24 / 60

GDPval turns

Astra / Fable

Those figures are workload observations, not a quote for your system. Measure prompt tokens, cached tokens, output tokens, retries, latency and human review on the workload you intend to operate.

The Artificial Analysis article describes the task-cost and token differences. Re-run the calculation when routing, caching, model versions or fallback policies change.

Capability is not containment

Safety evidence should sit beside capability evidence. A high capability rating does not prove that an agent is safe in your network, identity, data and tool environment.

OpenAI's deployment safety report illustrates why containment needs its own test record.

Anthropic's evaluation-incident report makes the same point from a separate provider context.

Ask what the agent can reach, what credentials it receives, what it can change, how actions are logged and who can stop it. Record the controls that were active during evaluation.

What the comparison means for an Australian assurance file

For Australian organisations, the model decision belongs inside the system's broader risk record. A model score cannot substitute for an owner, a use-case risk assessment, access controls, monitoring or an incident path.

APRA has warned that point-in-time assurance is poorly suited to systems that drift. Its AI letter supports a living evidence approach.

ASIC's 2026 cyber-risk communication makes the same governance direction relevant to financial services teams.

Link the ASIC release to the control owner and review cadence, not just the model card.

Minimum viable model-selection evidence

Create one record per decision. At minimum, include:

  • exact model and version;
  • benchmark version, date and source type;
  • harness, tools, permissions and reasoning setting;
  • fallback behaviour and traffic share;
  • cost, tokens, latency and retry data on your workload;
  • relevant safety, privacy and security results;
  • accountable owner and re-evaluation triggers.

How to choose if you must choose today

Choose Astra when your priority is efficient tool use, terminal work, document completion or a firm output-cost ceiling. Validate the result with your permissions and stopping rules.

Choose Fable when your priority is scientific, quantitative, analytical or long-context work. Validate the extra output and context cost against the value of the task.

If the decision is material, run both models on a representative, pre-registered workload. Include ordinary cases, edge cases, abstentions, tool errors and human review time.

Limits and closing principle

This comparison cannot establish training provenance, parameter access, Australian data residency, local-language performance or your organisation's containment quality. Those are separate evidence questions.

The defensible conclusion is narrow: published evidence can narrow the field, but it cannot make the decision for you. A model selection becomes assurance evidence only when the version, conditions, controls and owner are recorded together.

Primary sources