Which company has the third best AI model end of March?

Which company has the third best AI model end of March?

The hierarchy of large language models has shifted significantly over the last two weeks, creating a distinct “top tier” that seems to be hardening. With the recent release of OpenAI’s GPT-4o and Google’s major updates to Gemini 1.5 Pro, the competition for the top three spots on the LMSYS Chatbot Arena Leaderboard has become a game of inches. Here is the current state of play and why one specific company is currently locked into the third-place position.

Read more Bitcoin Up or Down — March 15, 9AM ET

Recent Developments and Fact-Check

The landscape changed on May 13, 2024, when OpenAI launched GPT-4o. This model immediately claimed the #1 spot on the leaderboard, pushing previous leaders down the ranking. Just a day later, on May 14, Google announced the wide availability of Gemini 1.5 Pro with enhanced reasoning capabilities, which saw a significant jump in its Arena score. Meanwhile, Anthropic has not released a major update to its Claude 3 family in the last 14 days, leading to a slight stagnation in its relative scoring as newer models flood the arena.

Currently, the leaderboard shows a fascinating trend: OpenAI often occupies both the first and second spots with different iterations of its models (such as GPT-4o and GPT-4 Turbo). This effectively leaves the “third best” slot as the primary battleground for competitors like Google and Anthropic.

The Case for Google

Google is the most likely candidate to hold the third-best model by the end of the period. Why? It comes down to the “OpenAI Sandwich.” OpenAI’s dominance is so pronounced that they frequently hold the top two positions on the leaderboard. This leaves Google’s flagship, Gemini 1.5 Pro, sitting firmly at number three.

Read more Bitcoin Up or Down — March 15, 6AM ET

Google has demonstrated a rapid-fire update cycle, moving Gemini 1.5 Pro from a preview stage to a production-ready model that rivals GPT-4’s reasoning. The model’s performance in the “Hard Prompts” and “Coding” categories has stabilized its score just below OpenAI’s peak but consistently above Anthropic’s Claude 3 Opus and xAI’s Grok-1.5. Unless Google manages to leapfrog OpenAI’s second-best model or falls behind a surprise release from a competitor, the third-place spot is its natural equilibrium.

The Competition: Anthropic and xAI

Anthropic’s Claude 3 Opus was a leader earlier this year, but without a “Claude 3.5” or “Claude 4” release in the immediate pipeline, it is losing ground to the sheer compute and iterative speed of Google. For Anthropic to take the third spot, they would need to either surpass Google (moving to #3) or see Google move up to #2, which is difficult given OpenAI’s dual-model strategy. As for xAI, while Grok-1.5 is a massive improvement over its predecessor, it still lacks the broad-based “vibes” and reasoning scores required to break into the top three of the Arena, which relies on human preference testing.

Key Triggers to Watch

What could change this picture? Keep an eye on these specific signals:

  • Anthropic Model Drops: Any announcement regarding “Claude 3.5” could immediately threaten Google’s third-place standing.
  • OpenAI Model Consolidation: If LMSYS decides to group OpenAI models differently, or if an older GPT-4 version falls out of favor, Google could inadvertently move up to #2.
  • DeepSeek Momentum: The DeepSeek-V2 model has been climbing the ranks; if its trajectory continues, it could disrupt the US-based dominance of the top 5.

Current data shows a very high conviction in Google’s position, with a 90% probability of them holding the third spot. Other contenders like Anthropic and xAI are trailing significantly, both hovering between 2% and 3% probability, reflecting the stability of the current leaderboard hierarchy.

Read more Bitcoin Up or Down on March 15?

Sources :

Leave a Reply

Your email address will not be published. Required fields are marked *