Zpět na Novinky
ProduktInstruktáž AI Understanding

Aplikace Pipette Liquid AI měří výkon umělé inteligence na zařízení na chytrých telefonech, uvádí BigGo

BigGo Finance uvádí, že Liquid AI vydala Pipette, bezplatnou aplikaci pro iOS a Android pro měření latence AI smartphonu, propustnosti a využití paměti. Zpráva dokumentuje test iPhonu 17 Pro a podrobně popisuje benchmark zaměřený na chytré telefony vyvinutý pomocí umělé analýzy.

5 min readRead the linked source
Source-provided image accompanying Liquid AI’s Pipette app measures on-device AI performance on smartphones, BigGo reports
Odkaz na zdrojZdroj zaznamenán
Vydavatel
finance.biggo.com
Odkaz na zdroj
finance.biggo.comhttps://finance.biggo.com/news/d74ec349-df7a-4f48-a861-4811784c123c
Typ zdroje
Propojený zdroj — stav primárního zdroje nebyl stanoven.
Také citováno

Příběh naposledy revidován

KontextPochopte to za 60 sekund

Začněte zde

Klíčové pojmy

Umělá inteligence na zařízení
Odvozování AI se provádí lokálně na uživatelském hardwaru, nikoli ve vzdálené cloudové službě.
Paměť (paměť agenta)
Uložený kontext, který agent AI používá v krocích nebo relacích ke zlepšení kontinuity.
Kvantování
Konverze modelových vah do formátů s nižší přesností, jako je 8bitový nebo 4bitový.
Otestujte seKvíz s vysvětlením modelů umělé inteligence

Co se změnilo od vydání

  1. Poprvé zveřejněno
  2. BigGo Finance materially advances the existing Pipette report with a hands-on iPhone 17 Pro workflow, a reported 48.3726-token-per-second Gemma 4 E2B IT result, an 18-minute test duration, details on the app’s four measurement modes and data-sharing option, and additional Artificial Analysis benchmark findings. The measurements and rankings are attributed to BigGo and are not independently confirmed here.

Co se stalo

Liquid AI released Pipette, a free smartphone app that lets users run quantized AI models locally and measure execution performance. BigGo Finance tested the iOS version on an iPhone 17 Pro and recorded 48.3726 tokens per second when running a 4-bit quantized Gemma 4 E2B IT model with MLX. BigGo also reported benchmark results from an Artificial Analysis collaboration covering small models that fit within 8GB when quantized.

BigGo Finance reports that Liquid AI released Pipette for iOS and Android as a free benchmark app for measuring AI model execution on smartphones. Users can download quantized models, including models from the Gemma and Qwen families, and run local tests for end-to-end latency, prefill throughput, decode throughput and maximum memory usage. The app requires account registration at first launch, according to BigGo, and downloaded models are stored on the device for testing. The source code is publicly available on GitHub, but this report does not independently verify the contents of that repository or the app’s availability in every market.

BigGo’s hands-on account describes a workflow in which users create a job, select either llama.cpp or MLX, choose a model and select test types. The report says all four test types are selected by default. It also says users can choose whether results are sent to the server as public data. Testing includes cooldown periods intended to reduce device temperature, and BigGo says its test of a 4-bit quantized Gemma 4 E2B IT model with MLX on an iPhone 17 Pro took 18 minutes. The resulting decode-throughput figure was 48.3726 tokens per second. These observations are reported by BigGo and have not been independently confirmed here.

BigGo also describes a separate collaboration between Artificial Analysis and Liquid AI to evaluate small models on an iPhone 17 Pro. The reported system combines Artificial Analysis’s intelligence-performance tests with Liquid AI’s inference-performance testing. Five benchmarks were used: BFCL, IFBench, AA-Omniscience, GPQA Diamond and MATH-500. BigGo says the evaluation was limited to models fitting within 8GB at 4-bit and used a context-length cap of 16,000 tokens for average-score comparisons. Nanbeige4.2-3B and LFM2.5-2.6B reportedly achieved the highest average scores under those conditions, while LFM2.5-8B-A1B led a comparison capped at one minute of response time. The source provides no independently checked raw benchmark table or replication results.

Podrobnosti o zdroji: finance.biggo.com ↗

Proč na tom záleží

Pipette makes practical constraints more visible by measuring latency, throughput and memory usage on the hardware where models actually run. That can help developers and users distinguish theoretical model quality from the speed and resource demands of local inference, although the results are platform-specific and not a universal ranking of models or phones.

The significance of Pipette is that it measures the conditions that determine whether local AI is usable on a particular device. A model may perform well on a conventional capability benchmark yet be too slow, memory-intensive or thermally demanding for a phone. Latency affects interactive use, throughput affects how quickly text can be generated, and memory use determines whether a model can run alongside the operating system and other applications. By exposing these measures, a tool such as Pipette can give developers more practical information when selecting models for offline or privacy-sensitive applications.

BigGo reports that conventional benchmarks are often designed around large, data-center-class models. According to the report, smaller models running on smartphones sometimes fail to complete those tests and receive scores of zero, while smartphone inference software is less mature and differs from data-center infrastructure. A benchmark tailored to small, quantized models may therefore provide a more useful picture of device-level tradeoffs. That does not make the results directly comparable with scores from larger systems. It means the evaluation is answering a narrower question: how capable and responsive a model is under specified local hardware and software conditions.

The reported results also illustrate why speed and intelligence cannot be treated as a single metric. BigGo says Nanbeige4.2-3B scored similarly to LFM2.5-2.6B but tended to require more time to produce 256 tokens. It also reports relatively low token efficiency for Qwen3.5 9B Reasoning and Qwen3.5 4B Reasoning in one analysis. These findings could matter for developers choosing between a more capable model and a faster one, but they should be treated as results from the reported test setup rather than general judgments about the models. No independent evaluation, user study or production deployment evidence is provided.

Interactive Mechanism

Interaktivní mechanismus: Jak to vlastně funguje

Interaktivně prozkoumejte základní technologii tohoto vývoje.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Interaktivní kontrola konceptu+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Na co se dále dívat

The main limitation is comparability. BigGo reports that iOS currently supports GPU acceleration while Android is CPU-only, and Liquid AI warns against cross-device comparisons. Future updates are expected to add models and comparisons based on memory usage, but the source does not independently establish the app’s broader reliability, adoption or long-term roadmap.

The first issue to watch is platform parity. BigGo reports that Pipette’s iOS version supports GPU-accelerated computation while the Android version is currently CPU-only. The report also identifies differences in flash-attention support, thread counts, accelerators and execution environments. Those differences can dominate performance results, so a score from an iPhone 17 Pro should not be read as a ranking of Android devices or as a direct comparison between operating systems. Liquid AI’s warning against cross-device comparisons is an important qualification.

The second issue is methodological transparency and reproducibility. Pipette can export results as CSV, according to BigGo, and its source code is publicly available. That could allow developers to inspect test configurations and repeat measurements. However, the source does not establish whether all model files, runtime versions, thermal conditions, compiler settings and benchmark implementations are identical across tests. It also does not report independent replication by researchers or users. Those details will determine how much confidence should be placed in differences of only a few tokens per second or small score changes.

Finally, the reported roadmap could broaden the tool’s usefulness if implemented as described. BigGo says Artificial Analysis plans to add models over time and compare models using memory usage rather than relying on one precision. That would help users evaluate tradeoffs among model size, quality, speed and resource consumption. Important unknowns remain: the source does not establish how widely Pipette is being used, whether the Android implementation will gain GPU support, how results are moderated or audited, or whether the benchmark predicts performance in real applications. Until those questions are answered, Pipette is best understood as a practical measurement tool for specific configurations, not a definitive leaderboard for smartphone AI.

Související průvodci a kvízy

Vysvětlení modelů AIŠkolení AITransformátoryOtestujte si, co víte – vyzkoušejte bezplatný kvíz AIVyhledejte si termín AI v našem slovníkuPostupujte podle sledování vydání modelu AI

Aktualizace a opravy

Tento kanonický příběh je aktualizován na místě, když se rozvíjející se událost podstatně změní. Jeho URL a původní datum vydání se nikdy nemění.

  • BigGo Finance materially advances the existing Pipette report with a hands-on iPhone 17 Pro workflow, a reported 48.3726-token-per-second Gemma 4 E2B IT result, an 18-minute test duration, details on the app’s four measurement modes and data-sharing option, and additional Artificial Analysis benchmark findings. The measurements and rankings are attributed to BigGo and are not independently confirmed here.
Viz veřejný protokol oprav
Považujete to za užitečné?