Ai showdown: google, openai, anthropic, xai battle for supremacy
The landscape of large language models (LLMs) is tightening. OpenLM.ai’s latest Chatbot Arena+ benchmark reveals a near-photographic finish between Google’s Gemini 3.1 Pro, OpenAI’s GPT 5.4, Anthropic’s Claude Opus 4.6, and xAI’s Grok 4.20, signaling a remarkable level of maturity in the field – and a potential end to the era of clear dominance.
The elo arena: more than just scores
Forget simplistic leaderboard rankings. The Chatbot Arena+ combines the rigorous Elo Arena system—powered by over 5 million human votes—with standardized technical metrics like AAII v3, MMLU-Pro, and ARC-AGI v2. This holistic approach provides a nuanced picture of performance, evaluating not just raw processing power, but also reasoning capabilities and subjective user preference. The ARC-AGI v2, in particular, highlights a persistent challenge: while humans achieve near-perfect scores on visual reasoning puzzles, even the most advanced AI models still hover between 10% and 20%.

Current top 5 llms: a march 2026 snapshot
Here’s a breakdown of the leading models, according to the latest data:
Gemini-3.1-Pro: Elo 1505, AAII v3 1310, MMLU-Pro 76%, ARC-AGI v2 91
Claude Opus 4.6 Thinking: Elo 1503, AAII v3 73, MMLU-Pro 89.7%, ARC-AGI v2 69.2
Grok-4.20: Elo 1496, AAII v3 72, MMLU-Pro 89.6%, ARC-AGI v2 38
GPT-5.4-high: Elo 1495, AAII v3 1290, MMLU-Pro 73%, ARC-AGI v2 88.5
Gemini-3-Pro: Elo 1492, AAII v3 1308, MMLU-Pro 73%, ARC-AGI v2 90

Subtle differences, distinct strategies
While the overall scores are remarkably close, the models differentiate themselves through specific strengths. Gemini 3.1 Pro excels in multimodal capabilities – seamlessly integrating text, image, and audio – alongside a balanced approach to logical reasoning and code generation. GPT 5.4 remains a powerhouse in programming and problem-solving, though its Elo score is slightly impacted by user preference for more “human-like” responses; a factor that triggered a user-driven push for OpenAI to reinstate earlier models like 4o.
Anthropic’s Claude 4.6 continues to prioritize safety and ethical considerations, making it a reliable choice for sensitive applications, while xAI’s Grok 4.20 is steadily gaining ground in conversational contexts. The ranking resonates with my own experiences, though I’d personally award Elon Musk's AI a slightly higher score in coding proficiency.

The chinese contenders falter
Despite the spotlight on Google, OpenAI, Anthropic, and xAI, the rankings reveal a surprising shift. Previously competitive Chinese models like GLM-4.6 (boasting a massive 200,000 token window) and Alibaba Cloud’s Qwen3.5-Max (over a trillion parameters) – once within striking distance of Gemini-2.5-Pro, Grok-4-0709, or GPT-5 – have noticeably dropped in the standings. This raises questions about the trajectory of Chinese AI development and its ability to compete at the highest level.

What this means for the future of ai
Gemini 3.1 Pro's current lead isn't definitive. The mere 30-point Elo gap between the top four demonstrates a remarkable level of maturity in LLMs. The resurgence of Chinese AI is also something to watch closely. Currently, the leading AI giants are carving out distinct niches: Google leads in multimodal integration, OpenAI maintains leadership in technical tasks and API compatibility, Anthropic champions safety and transparency, and xAI leans towards more emotionally resonant language. For users, this translates to an increasingly competitive market, with a wider range of options to suit specific needs—and the potential to leverage multiple models simultaneously.

Choosing the right tool for the job
Gemini Pro: Analyzing multimodal data (text + image). Ideal for auditing visual documents, scientific research.
GPT 5: Programming and algorithmic problem-solving. Essential for developers integrating with the Microsoft ecosystem.
Claude 4.5: Safety and code reliability. Suited for enterprise projects and secure environments.
Grok-4: Advanced conversational AI. Perfect for customer service and narrative analysis.

The price of power
Access to these cutting-edge models comes at a cost. Limited free tiers are available, with monthly subscription fees ranging from approximately €16 to €24, depending on the platform. With the gap between models shrinking, the emphasis is shifting from brute processing power to adaptability and integration within real-world ecosystems. The next openLM.ai ranking, expected in summer 2026, promises to offer a fresh perspective on evolving versions and the rising tide of open-source models.
The era of the dominant AI model is over. Now, it’s about who can best weave these powerful tools into the fabric of our daily lives, and that competition, it seems, is only just beginning.
n