Volver a Noticias
ProductoAI Understanding sesión informativa

Liquid AI’s Pipette app measures on-device AI performance on smartphones, BigGo reports

BigGo Finance reports that Liquid AI has released Pipette, a free iOS and Android app for measuring smartphone AI latency, throughput and memory use. The report documents an iPhone 17 Pro test and details a smartphone-focused benchmark developed with Artificial Analysis.

Por 5 min read
Editorial illustration of unbranded silicon processor modules, a small heat sink and a physical measurement scale representing latency, throughput and memory use; no identifiable people or readable text.
La versión corta

BigGo Finance reports that Liquid AI has released Pipette, a free iOS and Android app for measuring smartphone AI latency, throughput and memory use. The report documents an iPhone 17 Pro test and details a smartphone-focused benchmark developed with Artificial Analysis.

que paso

Liquid AI released Pipette, a free smartphone app that lets users run quantized AI models locally and measure execution performance. BigGo Finance tested the iOS version on an iPhone 17 Pro and recorded 48.3726 tokens per second when running a 4-bit quantized Gemma 4 E2B IT model with MLX. BigGo also reported benchmark results from an Artificial Analysis collaboration covering small models that fit within 8GB when quantized.

BigGo Finance reports that Liquid AI released Pipette for iOS and Android as a free benchmark app for measuring AI model execution on smartphones. Users can download quantized models, including models from the Gemma and Qwen families, and run local tests for end-to-end latency, prefill throughput, decode throughput and maximum memory usage. The app requires account registration at first launch, according to BigGo, and downloaded models are stored on the device for testing. The source code is publicly available on GitHub, but this report does not independently verify the contents of that repository or the app’s availability in every market.

BigGo’s hands-on account describes a workflow in which users create a job, select either llama.cpp or MLX, choose a model and select test types. The report says all four test types are selected by default. It also says users can choose whether results are sent to the server as public data. Testing includes cooldown periods intended to reduce device temperature, and BigGo says its test of a 4-bit quantized Gemma 4 E2B IT model with MLX on an iPhone 17 Pro took 18 minutes. The resulting decode-throughput figure was 48.3726 tokens per second. These observations are reported by BigGo and have not been independently confirmed here.

BigGo also describes a separate collaboration between Artificial Analysis and Liquid AI to evaluate small models on an iPhone 17 Pro. The reported system combines Artificial Analysis’s intelligence-performance tests with Liquid AI’s inference-performance testing. Five benchmarks were used: BFCL, IFBench, AA-Omniscience, GPQA Diamond and MATH-500. BigGo says the evaluation was limited to models fitting within 8GB at 4-bit quantization and used a context-length cap of 16,000 tokens for average-score comparisons. Nanbeige4.2-3B and LFM2.5-2.6B reportedly achieved the highest average scores under those conditions, while LFM2.5-8B-A1B led a comparison capped at one minute of response time. The source provides no independently checked raw benchmark table or replication results.

Lea la fuente principal: finance.biggo.com

Por qué es importante

Pipette makes practical on-device AI constraints more visible by measuring latency, throughput and memory usage on the hardware where models actually run. That can help developers and users distinguish theoretical model quality from the speed and resource demands of local inference, although the results are platform-specific and not a universal ranking of models or phones.

The significance of Pipette is that it measures the conditions that determine whether local AI is usable on a particular device. A model may perform well on a conventional capability benchmark yet be too slow, memory-intensive or thermally demanding for a phone. Latency affects interactive use, throughput affects how quickly text can be generated, and memory use determines whether a model can run alongside the operating system and other applications. By exposing these measures, a tool such as Pipette can give developers more practical information when selecting models for offline or privacy-sensitive applications.

BigGo reports that conventional benchmarks are often designed around large, data-center-class models. According to the report, smaller models running on smartphones sometimes fail to complete those tests and receive scores of zero, while smartphone inference software is less mature and differs from data-center infrastructure. A benchmark tailored to small, quantized models may therefore provide a more useful picture of device-level tradeoffs. That does not make the results directly comparable with scores from larger systems. It means the evaluation is answering a narrower question: how capable and responsive a model is under specified local hardware and software conditions.

The reported results also illustrate why speed and intelligence cannot be treated as a single metric. BigGo says Nanbeige4.2-3B scored similarly to LFM2.5-2.6B but tended to require more time to produce 256 tokens. It also reports relatively low token efficiency for Qwen3.5 9B Reasoning and Qwen3.5 4B Reasoning in one analysis. These findings could matter for developers choosing between a more capable model and a faster one, but they should be treated as results from the reported test setup rather than general judgments about the models. No independent evaluation, user study or production deployment evidence is provided.

Qué ver a continuación

The main limitation is comparability. BigGo reports that iOS currently supports GPU acceleration while Android is CPU-only, and Liquid AI warns against cross-device comparisons. Future updates are expected to add models and comparisons based on memory usage, but the source does not independently establish the app’s broader reliability, adoption or long-term roadmap.

The first issue to watch is platform parity. BigGo reports that Pipette’s iOS version supports GPU-accelerated computation while the Android version is currently CPU-only. The report also identifies differences in flash-attention support, thread counts, accelerators and execution environments. Those differences can dominate performance results, so a score from an iPhone 17 Pro should not be read as a ranking of Android devices or as a direct comparison between operating systems. Liquid AI’s warning against cross-device comparisons is an important qualification.

The second issue is methodological transparency and reproducibility. Pipette can export results as CSV, according to BigGo, and its source code is publicly available. That could allow developers to inspect test configurations and repeat measurements. However, the source does not establish whether all model files, runtime versions, thermal conditions, compiler settings and benchmark implementations are identical across tests. It also does not report independent replication by researchers or users. Those details will determine how much confidence should be placed in differences of only a few tokens per second or small score changes.

Finally, the reported roadmap could broaden the tool’s usefulness if implemented as described. BigGo says Artificial Analysis plans to add models over time and compare models using memory usage rather than relying on one quantization precision. That would help users evaluate tradeoffs among model size, quality, speed and resource consumption. Important unknowns remain: the source does not establish how widely Pipette is being used, whether the Android implementation will gain GPU support, how results are moderated or audited, or whether the benchmark predicts performance in real applications. Until those questions are answered, Pipette is best understood as a practical measurement tool for specific configurations, not a definitive leaderboard for smartphone AI.

Guías y cuestionarios relacionados

Modelos de IA explicadosEntrenamiento de IAtransformadoresPon a prueba lo que sabes: prueba un cuestionario gratuito sobre IABusque un término de IA en nuestro glosario
¿Encontró esto útil?