What is GPT-6 Provider Adapter

2026-09-05

In my last post I wrote Astra hit 99.9% on ARC-AGI-3 and called the benchmark dead. But the leaderboard shows two numbers: 62.7% and 99.9%.

The 37% gap is the harness.

When ARC-AGI-3 launched 6 months ago, the best frontier model scored 0.51%. François Chollet predicted it would take a year to crack. Astra did it in 5 months, 2x faster than expected.

Standard Harness

manual_rolling
Prompt
Orders model to write and carry forward manual text notes.
Output
Action plus manual ASCII scratchpad.
L8 hub q2(8v) rot=49,18 ext8 to3; ret10 to2
State
Internal reasoning tokens dropped after every move.
Turn N+1
Old messages trimmed. Model reconstructs state from text notes.
Score: 62.7%
Cost: $26,098
3.66x slower

Provider Adapter

continuous_conversation
Prompt
Note-taking prompt automatically stripped. Zero scratchpad bloat.
Output
Pure discrete action without text notes.
ACTION6 39 4
State
Encrypted reasoning tokens chained directly to next turn.
Turn N+1
175k compaction. Living working memory stays intact.
Score: 99.9%
Cost: $18,817
49% fewer tokens

The Standard Harness (62.7%)

ARC Prize's default setup is the Standard harness (manual_rolling). Provider-neutral, 64x64 text grid, turn by turn.

It wipes internal reasoning tokens after every move. The prompt forces the model to keep manual text notes: Carry forward notes you choose to keep with it throughout the environment.

Astra burns output tokens writing ASCII scratchpads every turn:

When the conversation gets long, older turns get trimmed. If a hallucinated rule gets typed into notes, it poisons every turn after.

Scores:

What is the Provider Adapter? (99.9%)

The Provider Adapter harness (continuous_conversation) drops the scratchpad entirely.

Uses OpenAI's Responses API:

  1. Latent reasoning persistence: Runs stateless (store: false, ZDR-compatible). Catches reasoning.encrypted_content and passes the encrypted blob back into turn N+1 as raw input items.
  2. Trained model compaction: Compaction isn't an external script. OpenAI trained Astra itself to analyze conversation history and emit an encrypted compaction item at 175k tokens (compact_threshold).
  3. No note-taking prompt: Deletes the manual carry-forward instruction. Astra outputs pure actions.

Beat human action baseline on 96% of levels (-51.7% actions/level).

The Cost Inversion

Higher reasoning effort is cheaper on ARC-AGI-3:

Smart moves solve games in fewer actions. Fewer actions = fewer API calls = cheaper bill. Provider Adapter at Max reasoning is 65% cheaper than Standard harness at None.

The Others: PRO-LONG, Tycho, VISTA

NVIDIA AVO Hits 100%

NVIDIA posted their own ARC-AGI-3 run right before Astra.

100.00 RHAE on all 25 public environments. All 183 levels cleared using Opus 5.

Official ARC board has Claude Opus 5 at 30.2%. 100% in NVIDIA's harness instead.

NVIDIA used AVO (Agentic Variation Operators), originally built to tune GPU kernels on DGX B200s. Ran 7 days straight on attention kernels, tested 500+ directions, beat cuDNN by +3.5% and FlashAttention-4 by +10.5%.

Plugged into ARC-AGI-3 with zero changes:

6,624 actions across all 183 levels. 12% fewer actions than VISTA with the same model.

Thoughts

Thinking of building a coding harness to test if any of these findings actually transfer over. Will update soon.

written by me reviewed by ai

Visitors: -