A 438-billion-parameter model designed for coding and enterprise agents reached a score of 43 on Artificial Analysis' Intelligence Index while delivering approximately 183 tokens per second, according to Multiverse Computing's Wednesday launch of Quasar 438B. The Spanish company is positioning Quasar as the highest-scoring European model on the Intelligence Index, surpassing Mistral Medium 3.5 and NVIDIA Nemotron 3 Ultra. Multiverse built the system using compression technology intended to make large-scale reasoning models fast and cheap enough for agents that repeatedly reason, call tools, and verify results.
Quasar achieved 69.3 on Terminal-Bench v2.1, placing it ahead of Mistral Medium 3.5 but significantly behind top frontier systems, which Claude Opus 5 leads at 89.1. Artificial Analysis recorded Quasar's time to first token at roughly 1.1 seconds, with a complete 500-token response including reasoning delivered in approximately 15.3 seconds. The model features a one-million-token context window and supports English and Spanish through the Multiverse CompactifAI API. Multiverse raised $570 million in a July Series C round to expand its library of compressed models and commercialize the underlying technology.
The company claims its CompactifAI system can shrink large AI models by 80% to 95% with minimal accuracy loss, though it hasn't disclosed how much Quasar was compressed or which base model it started from. Multiverse hasn't revealed what hardware Quasar requires to operate or how substantially the compression reduces memory and compute demands. The report notes that these details matter for agents, which may call the model and other tools repeatedly before completing a task. Quasar remains proprietary and accessible only through Multiverse's API, preventing developers from inspecting weights or running it on their own infrastructure.
Multiverse is targeting software engineering, technical copilots, research, and workflow automation as primary use cases for Quasar. The million-token context window provides agents with space to handle large codebases and retain information throughout task execution, though processing additional context demands more compute. The report explains that agent latency extends beyond raw throughput—agents must wait for tools, handle expanding context, and execute multiple model calls during a single task, while the agent tooling layer itself continues developing to meet model requirements. Quasar arrives as European AI companies increasingly build their own model and compute infrastructure rather than depending on U.S. hyperscalers, though Multiverse is pursuing compression as an alternative path to making massive models economically viable. Compression as a strategy for deployment represents a fundamental bet on architecture efficiency over raw parameter scaling, and whether that trade-off proves commercially durable will likely depend on use cases where repeated inference cost matters more than peak capability.

