In January 2026, VentureBeat reported that Artificial Analysis released a major overhaul of its AI Intelligence Index. The firm shifted emphasis toward tests it describes as more real-world: agents, coding, scientific reasoning, and general knowledge with equal weighting in the aggregate score.
How to use an index without over-trusting it
- Read category scores, not only the headline rank. Equal weighting can hide strengths and weaknesses.
- Use public indices for orientation only.
- Require vendor evals with published prompts when possible.
- Run a private eval set of 50 to 200 tasks from your workflows, graded by humans.
- Record the eval-set version and date on your shortlist.
When open-weight models post strong coding numbers under permissive licenses (for example VentureBeat’s GLM-5.1 coverage), re-run tasks you care about. License terms, hosting region, and safety refusal behavior can dominate raw scores for regulated industries.





