Gemma 4 by Google DeepMind: five open-weight models (E2B to 31B Dense) under Apache 2.0. CI-First Benefit Score 7.0 (Strong). Arena AI #3 open model. Free to self-host with frontier-level reasoning, multimodal input, and agentic function-calling.
Mistral Large 3 earns a CI-First Benefit Score of 7.0/10 and a Humics-Neutral badge. Its 675B-parameter sparse architecture, 41B active parameters, 256K context window, multimodal input, and Apache 2.0 weights support demanding research, coding, multilingual, and enterprise workflows. Strong human review remains necessary because fluent output can still contain factual, analytical, and code errors.
Grok 4.6: CI-First Benefit Score 6.8/10, Humics-Neutral. xAI's frontier model for coding, agentic tasks, and knowledge work with a 500K context window, configurable reasoning, and real-time web search. Ranked #6 on the Artificial Analysis Intelligence Index with a score of 61.