The artificial intelligence landscape is currently fixated on a single, daunting hurdle: “Humanity’s Last Exam” (HLE). Developed by Scale AI and the Center for AI Safety, this benchmark isn’t your typical high school level test. It consists of over 5,000 questions across dozens of specialized fields, designed specifically to be so difficult that even subject-matter experts struggle. As Google prepares its next-generation Gemini 3 architecture, the question isn’t just whether it will pass, but how deep into the “expert” territory it can climb before the mid-2026 deadline.
Read more Bitcoin Up or Down — April 10, 11AM ET
Recent Developments and Context
In the last few weeks, the trajectory for high-level reasoning models has become clearer. Google’s release of Gemini 2.0 in December 2024 signaled a shift from general multimodal capabilities toward specialized “thinking” and reasoning modes. Here is what is currently shaping the outlook for Gemini 3:
- The Reasoning Pivot: Google DeepMind has integrated advanced search and “thinking” capabilities into its latest iterations, directly competing with OpenAI’s o1 and o3 models. These models are specifically designed to handle the multi-step logic required by the HLE.
- Benchmark Difficulty: Current state-of-the-art models are generally hovering in the 30% to 40% range on the HLE. For a model to hit the 45% or 50% mark, it requires more than just a larger dataset; it needs a fundamental leap in how it verifies its own logic.
- Compute Scaling: Google’s internal roadmap suggests that Gemini 3 will leverage significantly more TPU (Tensor Processing Unit) resources than its predecessor, aiming for a generational jump in “system 2” thinking—the slow, deliberate reasoning necessary for expert-level science and math.
The Case for the 45% Threshold
The most grounded expectation is that Gemini 3 will comfortably clear the 45% mark. Why? Because the jump from the current 2.0 performance to 3.0 is expected to incorporate “test-time compute” scaling. This technique allows the model to “think” longer before answering, which has historically led to massive gains in difficult reasoning benchmarks. Given that current top-tier models are already knocking on the door of 40%, a next-generation release with a year of additional optimization makes 45% look like a baseline rather than a ceiling.
Comparing the Competitors: 50% vs. 55%
While 45% seems highly probable, the 50% and 55% thresholds are where the real debate lies. A 50% score would represent a landmark achievement, effectively placing the AI at a level of proficiency comparable to a PhD-level expert across every single subject in the exam. While Gemini 3 is expected to be a powerhouse, the HLE is specifically designed to resist “gaming” via training data. Achieving 55% or 60% would require the model to solve problems that currently baffle most human experts, making those higher targets significantly more speculative at this stage of development.
Read more What price will Bitcoin hit on April 10?
What to Watch For
Several triggers will likely shift these expectations in the coming months. First, keep an eye on the release of “Gemini 2.0 Thinking” or “Pro” versions; their performance on the HLE leaderboard will serve as the definitive floor for what Gemini 3 will eventually achieve. Second, any updates to the HLE leaderboard itself—such as new models from competitors hitting the 45% mark—will validate that the “reasoning wall” is being broken.
Current data shows near-certainty (over 99%) for the 40% and 45% milestones, reflecting a consensus that the next generation of hardware and software will inevitably surpass these levels. However, the outlook for the 50% mark is more divided, currently sitting at approximately 38.5% probability, while the 55% and 60% targets remain low-confidence outliers at 18% and 5% respectively. Liquidity remains highest for the 40% and 45% segments, totaling over $50,000 in combined volume.
Read more US x Iran permanent peace deal by…?
Sources :