Gemini 4 Argon by @GoogleDeepMind just went head-to-head with top frontier models across a set of generations. At 33% lower cost per task than GPT-6.1 Sol, and 70% lower than Claude Opus 5.5, see first impressions from @petergostev on how Gemini 4 Argon compares.

Arena.ai @arena ·
  1. #1

    This Week in the Arena: Four releases reshaped the Pareto frontier across Text, Code, and Agent Arena. - After OpenAI’s DevDay, GPT-6.1 Sol (Max) entered Code Arena: WebDev at #3…

  2. #2

    How to design rewards for post-training frontier image models? Our research suggests human preference reward is necessary, but insufficient: A preference model may still reward ou…

  3. #3

    Arena Open House: Rooftop Happy Hour 10/7. We're throwing open the doors to our new SF HQ office and want to invite researchers, developers, and builders pushing on hard AI proble…

  4. #4

    Introducing Agent Mode: Agentic AI is now measured in the Arena. Agent Mode can do deep research, create reports, generate images, build websites, debug code, and more. It complet…

查看 @arena 的全部帖子