A U.S.-based artificial intelligence firm, Liquid AI, has introduced a groundbreaking model that operates entirely on smartphones and other edge computing devices. The newly released LFM2.5-2.6B, featuring 2.69 billion parameters, is now accessible for free download from platforms such as Hugging Face. This development marks a significant shift in how AI agents are deployed and used.
Traditionally, AI agents required access to cloud-based infrastructure to perform complex tasks. The LFM2.5-2.6B overturns this norm by enabling complete local processing on individual devices. This change significantly cuts down on latency, enhances data privacy, and eliminates cloud-related costs that had become the industry standard.
Model Performance and Capabilities
The LFM2.5-2.6B is trained on a massive dataset containing 34 trillion tokens, giving it a robust foundation for performance. It supports a context length of 131,072 tokens and includes a vocabulary of 128,000 words. The model comes in multiple formats, including its native form, GGUF for the open-source llama.cpp framework, ONNX for cross-platform operations, and MLX for Apple's silicon chips. Developers can choose between a base version suitable for fine-tuning or a specialized version optimized for agent tasks.
On the BFCLv4 benchmark, which evaluates tool-calling capability, the LFM2.5-2.6B scored 56.88. This result surpasses Google’s 5.18B parameter Gemma 4 E28 at 36.98 and the 4.7B parameter Qwen3.5 at 50.56. It nearly matches the 60.13 score of the 9.7B parameter Qwen3.5-9B model. While the compared models are multimodal, the LFM2.5-2.6B’s performance suggests strong single-modal capabilities.
Speed is another major focus of the model. On Apple's M5 Max, it processes 220 tokens per second. On AMD's Ryzen AI Max+ 395, it handles 113 tokens per second. On smartphones, it delivers 30 tokens per second while keeping memory consumption under 2.5GB. With a GPU like the NVIDIA H100, it can generate 15,000 tokens per second under high concurrency, translating to over 1.3 billion tokens per day.
Economic and Market Implications
The current AI landscape is marked by an intense price war for cloud-based inference services. OpenAI’s GPT-5.6 charges $2 per million input tokens and $6 per million output tokens, while DeepSeek’s V4-Flash model costs $0.14 for input and $0.28 for output. The release of the LFM2.5-2.6B at no cost threatens this model by offering high performance without reliance on paid cloud services.
Liquid AI has made the model available free of charge to developers, researchers, and commercial users with annual revenues under $10 million (around ¥1.5 billion). This pricing strategy positions the LFM2.5-2.6B as a zero-cost alternative that even DeepSeek's aggressive pricing cannot match. The move could push cloud service providers toward higher-margin work, such as training large-scale models where computational resources are more constrained.
A recent $10 billion deal between data center investor Volta and AI firm Anthropic highlights this shift. The partnership is a sign that training compute will remain a scarce and valuable resource. Meanwhile, the LFM2.5-2.6B's lightweight design makes it ideal for robotics and other physical systems that require local decision-making. With a context window of 128,000 tokens and compact memory use, it could serve as the core intelligence for next-generation edge robotics.
The model’s arrival could redefine the AI economy by decentralizing processing. Future AI systems may see large cloud models handling strategic planning, while smaller, device-based models manage real-time execution. This dual-track approach could lead to faster, more secure, and more efficient AI applications in both consumer and industrial contexts.

