Evidence · DiggingBeagle record
The Range Shrinks, the Threat Remains: Re-evaluating LLM Package Hallucinations on the 2026 Frontier-Model Cohort
A 2026 arXiv preprint replicates earlier package-hallucination testing on Claude Sonnet 4.6, Claude Haiku 4.5, GPT-5.4-mini, Gemini 2.5 Pro, and DeepSeek V3.2. Across 199,845 paired Python and JavaScript prompts, it reports overall hallucination rates from 4.62% to 6.10% and identifies 127 package names invented identically by all five tested models. The work is a preprint and should not be treated as a clean longitudinal comparison with the 2025 USENIX study because the model cohort and methodology differ.
- Published
- May 16, 2026
- Publisher
- arXiv
Evidence record
A 2026 arXiv preprint replicates earlier package-hallucination testing on Claude Sonnet 4.6, Claude Haiku 4.5, GPT-5.4-mini, Gemini 2.5 Pro, and DeepSeek V3.2. Across 199,845 paired Python and JavaScript prompts, it reports overall hallucination rates from 4.62% to 6.10% and identifies 127 package names invented identically by all five tested models. The work is a preprint and should not be treated as a clean longitudinal comparison with the 2025 USENIX study because the model cohort and methodology differ.