Grok 4.7 is first in the new legal AI test with 9.4%
Artificial Analysis says a hallucination is material if it can mislead a reader, for example with an incorrect contract date.
Artificial Analysis and Harvey announced version 1.1 of the score method for the Legal Agent Benchmark. The new method adds a hallucination check. A task counts only if the work follows all the rubric criteria and has no material hallucination. Grok 4.7 from SpaceXAI is first with 9.4%. Artificial Analysis says that more than 60% of the results that are correct without the check contain a material hallucination.
Sources
Posted