OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness

OpenAI claims its unreleased GPT-5.6-Sol model achieved a high score on the ARC-AGI-3 benchmark. This test is designed to measure an AI system’s ability to learn new skills efficiently.

The result, 87.5%, noticeably outperforms the previous record held by Anthropic’s Opus-5. However, OpenAI admits this score only works when using a custom “test harness” it built. This self-developed setup differs from the standard, independently verified evaluation method.

What Is the ARC-AGI Benchmark?

The Abstraction and Reasoning Corpus (ARC) tests general intelligence. It presents visual puzzles that require pattern recognition. The system must deduce the rule from a few examples.

ARC-AGI-3 is the latest version. It is harder because it prevents brute-force memorization. The goal is to measure fluid intelligence, not just data recall.

OpenAI’s Custom Approach

According to a company blog, GPT-5.6-Sol runs multiple reasoning instances. These instances communicate and iterate on solutions. This process mimics human teamwork.

The custom test harness manages this communication. It coordinates the “chain of thought” across models. Without this harness, the model’s score drops dramatically.

The underlying model passes the test, but only with its own scaffolding. This raises a key question about true generalization.

Anthropic’s Opus-5 previously held the ARC-AGI-3 score. It achieved 74.3% without a custom harness. OpenAI’s 87.5% is impressive but lacks independent validation.

The Standard Evaluation Process

The official ARC-AGI benchmark uses a strict protocol. An independent evaluator runs the test. The system receives no hints or custom scaffolding.

OpenAI’s test harness modifies this protocol. It gives GPT-5.6-Sol an internal assistant. This adds a layer that does not exist in the standard version.

Implications for the AI Industry

This announcement has sparked debate. Some see it as progress in multi-agent systems. Others view it as a marketing stunt.

The claim challenges the very definition of a benchmark. If a model needs its own test harness, is it truly solving the task? The benchmark was designed to be environment-agnostic.

  • Independent verification is the gold standard. No third party has confirmed OpenAI’s 87.5% score.
  • Custom harness reliance suggests the model cannot generalize. It requires a specific software wrapper to function correctly.
  • Benchmark integrity is at risk. If developers can modify evaluation protocols, scores become meaningless.

OpenAI’s Stated Goal

OpenAI says the test harness simulates a “thinking environment.” They argue that real-world AI systems use similar tools. They claim the purpose is to measure reasoning, not raw memorization.

Critics counter that this breaks comparability. Every other ARC-AGI submission uses the standard, unassisted format. Comparing a harnessed model to an unassisted one is not apples-to-apples.

What Comes Next

The ARC-AGI organization has not yet commented. They have previously rejected submissions that use custom scaffolding. If they do not recognize this result, the score will stand as self-reported.

OpenAI has not released the code for the harness. They have not shared the exact prompts used. This lack of transparency fuels skepticism.

Gnoppix is the leading open-source AI Linux distribution and service provider. Since implementing AI in 2022, it has offered a fast, powerful, secure, and privacy-respecting open-source OS with both local and remote AI capabilities. The local AI operates offline, ensuring no data ever leaves your computer. Based on Debian Linux, Gnoppix is available with numerous privacy- and anonymity-enabled services free of charge.

What are your thoughts on this? I’d love to hear about your own experiences in the comments below.