Claude Fable 5.1 held on to the number one spot on Artificial Analysis’ Intelligence Index v4.2, and tied with OpenAI’s new model on the updated benchmarks 🏆. What you need to know: 📊 Scores 57 on the v4.2 index, edging out GPT 6 Astra. 🔬 Achieves 52.6% on Terminal Bench

  1. #1

    This assistant remembers where you left your keys. No network connection. No remote server. Build one yourself in a course built in partnership with Qdrant and taught by Dylan Cou…

  2. #2

    A model trained only to pass tests learned to add unrequested code and let errors pass silently. 🧪 Xiaomi fixed it by multiplying each test result by quality checklist scores. Res…

  3. #3

    An open model nearly matches Claude Mythos on cybersecurity capabilities: GLM-5.3 solves 12% and Claude Mythos 14% of ExploitBench exploit tasks, per Anthropic. Just $20.40 in tok…

  4. #4

    This week in The Batch, Andrew's Letter covers how open weights model GLM-5.3 nearly matched Claude Mythos at exploiting vulnerabilities, 12% vs. 14%. Also inside: 🧪 Xiaomi's new…

  5. #5

    Andrew Ng's AI Engineering Skills Map shows what to learn. AI Dev is where you hear from the people who built it in the real world. Keynotes from Andrew Ng and Yann LeCun. Enginee…

  6. #6

    The Data Engineering Professional Certificate is now available on DeepLearningAI. Across four courses, design and build the systems that generate, ingest, store, transform, and se…

查看 @DeepLearningAI 的全部帖子