In my last post I wrote Astra hit 99.9% on ARC-AGI-3 and called the benchmark dead. But the leaderboard shows two numbers: 62.7% and 99.9%.
The 37% gap is the harness.
When ARC-AGI-3 launched 6 months ago, the best frontier model scored 0.51%. François Chollet predicted it would take a year to crack. Astra did it in 5 months, 2x faster than expected.
Standard Harness
Provider Adapter
The Standard Harness (62.7%)
ARC Prize's default setup is the Standard harness (manual_rolling). Provider-neutral, 64x64 text grid, turn by turn.
It wipes internal reasoning tokens after every move. The prompt forces the model to keep manual text notes: Carry forward notes you choose to keep with it throughout the environment.
Astra burns output tokens writing ASCII scratchpads every turn:
- Coordinates:
9−=(39,4), rotate=(49,18), 14+=(59,11) - State:
Turn 5: P=(24,20), empty, facing west - Ad-hoc rules:
extend8 to3; retract10 to2; shorten8 to1
When the conversation gets long, older turns get trimmed. If a hallucinated rule gets typed into notes, it poisons every turn after.
Scores:
- Max reasoning: 62.7% ($26,098)
- High reasoning: 54.8% ($40,705)
- None (zero thinking): 35.2% ($49,791)
What is the Provider Adapter? (99.9%)
The Provider Adapter harness (continuous_conversation) drops the scratchpad entirely.
Uses OpenAI's Responses API:
- Latent reasoning persistence: Runs stateless (
store: false, ZDR-compatible). Catchesreasoning.encrypted_contentand passes the encrypted blob back into turn N+1 as raw input items. - Trained model compaction: Compaction isn't an external script. OpenAI trained Astra itself to analyze conversation history and emit an encrypted compaction item at 175k tokens (
compact_threshold). - No note-taking prompt: Deletes the manual carry-forward instruction. Astra outputs pure actions.
- High reasoning: 99.9% ($18,817)
- Max reasoning: 98.6% ($17,332)
- None (zero thinking): 96.7% ($23,457)
Zero per-turn reasoning tokens, yet 96.7%. The accumulated reasoning state from prior turns is already warm in cache. It's 3.66x faster in wall-clock time and uses 49% fewer total tokens.
Beat human action baseline on 96% of levels (-51.7% actions/level).
The Cost Inversion
Higher reasoning effort is cheaper on ARC-AGI-3:
- Astra Standard (None): $49,791
- Astra Standard (Max): $26,098
- Astra Provider Adapter (None): $23,457
- Astra Provider Adapter (Max): $17,332
Smart moves solve games in fewer actions. Fewer actions = fewer API calls = cheaper bill. Provider Adapter at Max reasoning is 65% cheaper than Standard harness at None.
The Others: PRO-LONG, Tycho, VISTA
- PRO-LONG stored interaction logs on disk and searched them with code tools instead of prompt sliding windows. 4.2x to 5.8x fewer tokens. Fable 5 hit 97.4% for just $1,750.
- Tycho had an agent write an executable Python simulator per game. Sol and Opus 5 both hit 100.00 RHAE. But auto-repairing simulator bugs dropped efficiency (88.5 down to 83.1). Simulating transitions accurately is not the same as solving the objective.
- VISTA hooked Claude Opus 5 to Claude Code with rendered 512x512 PNG frames. Cleared the 183 public levels in 7,542 actions.
NVIDIA AVO Hits 100%
NVIDIA posted their own ARC-AGI-3 run right before Astra.
100.00 RHAE on all 25 public environments. All 183 levels cleared using Opus 5.
Official ARC board has Claude Opus 5 at 30.2%. 100% in NVIDIA's harness instead.
NVIDIA used AVO (Agentic Variation Operators), originally built to tune GPU kernels on DGX B200s. Ran 7 days straight on attention kernels, tested 500+ directions, beat cuDNN by +3.5% and FlashAttention-4 by +10.5%.
Plugged into ARC-AGI-3 with zero changes:
- 64x64 raw text grid: Zero vision tokens.
- Persistent memory: Keeps prior attempts, error traces, and hypotheses alive across resets.
- Supervisor agent: Watches the trajectory. If the worker agent gets stuck in a loop or plateaus, the supervisor interrupts and redirects it.
6,624 actions across all 183 levels. 12% fewer actions than VISTA with the same model.
Thoughts
Thinking of building a coding harness to test if any of these findings actually transfer over. Will update soon.