Quay lại Tin tức
sản phẩmAI Understanding tóm tắt

Ứng dụng Pipette của Liquid AI đo lường hiệu suất AI trên thiết bị trên điện thoại thông minh, báo cáo của BigGo

BigGo Finance báo cáo rằng Liquid AI đã phát hành Pipette, một ứng dụng iOS và Android miễn phí để đo độ trễ, thông lượng và mức sử dụng bộ nhớ AI trên điện thoại thông minh. Báo cáo ghi lại bài kiểm tra iPhone 17 Pro và nêu chi tiết điểm chuẩn tập trung vào điện thoại thông minh được phát triển bằng Phân tích nhân tạo.

5 min readRead the linked source
Source-provided image accompanying Liquid AI’s Pipette app measures on-device AI performance on smartphones, BigGo reports
Nguồn tham khảoNguồn đã ghi
Nhà xuất bản
finance.biggo.com
Liên kết nguồn
finance.biggo.comhttps://finance.biggo.com/news/d74ec349-df7a-4f48-a861-4811784c123c
Loại nguồn
Nguồn được liên kết - trạng thái nguồn chính chưa được thiết lập.
Cũng được trích dẫn

Câu chuyện được sửa đổi lần cuối

Bối cảnhHiểu điều này trong 60 giây

Bắt đầu ở đây

Thuật ngữ chính

AI trên thiết bị
Suy luận AI được thực hiện cục bộ trên phần cứng của người dùng thay vì trên dịch vụ đám mây từ xa.
Bộ nhớ (Bộ nhớ tác nhân)
Bối cảnh được lưu trữ mà tác nhân AI sử dụng qua các bước hoặc phiên để cải thiện tính liên tục.
Lượng tử hóa
Chuyển đổi trọng số mô hình sang các định dạng có độ chính xác thấp hơn như 8 bit hoặc 4 bit.
Tự kiểm traCâu đố giải thích về mô hình AI

Điều gì đã thay đổi kể từ khi xuất bản

  1. Xuất bản lần đầu
  2. BigGo Finance materially advances the existing Pipette report with a hands-on iPhone 17 Pro workflow, a reported 48.3726-token-per-second Gemma 4 E2B IT result, an 18-minute test duration, details on the app’s four measurement modes and data-sharing option, and additional Artificial Analysis benchmark findings. The measurements and rankings are attributed to BigGo and are not independently confirmed here.

Chuyện gì đã xảy ra

Liquid AI released Pipette, a free smartphone app that lets users run quantized AI models locally and measure execution performance. BigGo Finance tested the iOS version on an iPhone 17 Pro and recorded 48.3726 tokens per second when running a 4-bit quantized Gemma 4 E2B IT model with MLX. BigGo also reported benchmark results from an Artificial Analysis collaboration covering small models that fit within 8GB when quantized.

BigGo Finance reports that Liquid AI released Pipette for iOS and Android as a free benchmark app for measuring AI model execution on smartphones. Users can download quantized models, including models from the Gemma and Qwen families, and run local tests for end-to-end latency, prefill throughput, decode throughput and maximum memory usage. The app requires account registration at first launch, according to BigGo, and downloaded models are stored on the device for testing. The source code is publicly available on GitHub, but this report does not independently verify the contents of that repository or the app’s availability in every market.

BigGo’s hands-on account describes a workflow in which users create a job, select either llama.cpp or MLX, choose a model and select test types. The report says all four test types are selected by default. It also says users can choose whether results are sent to the server as public data. Testing includes cooldown periods intended to reduce device temperature, and BigGo says its test of a 4-bit quantized Gemma 4 E2B IT model with MLX on an iPhone 17 Pro took 18 minutes. The resulting decode-throughput figure was 48.3726 tokens per second. These observations are reported by BigGo and have not been independently confirmed here.

BigGo also describes a separate collaboration between Artificial Analysis and Liquid AI to evaluate small models on an iPhone 17 Pro. The reported system combines Artificial Analysis’s intelligence-performance tests with Liquid AI’s inference-performance testing. Five benchmarks were used: BFCL, IFBench, AA-Omniscience, GPQA Diamond and MATH-500. BigGo says the evaluation was limited to models fitting within 8GB at 4-bit and used a context-length cap of 16,000 tokens for average-score comparisons. Nanbeige4.2-3B and LFM2.5-2.6B reportedly achieved the highest average scores under those conditions, while LFM2.5-8B-A1B led a comparison capped at one minute of response time. The source provides no independently checked raw benchmark table or replication results.

Chi tiết nguồn: finance.biggo.com ↗

Tại sao nó quan trọng

Pipette makes practical constraints more visible by measuring latency, throughput and memory usage on the hardware where models actually run. That can help developers and users distinguish theoretical model quality from the speed and resource demands of local inference, although the results are platform-specific and not a universal ranking of models or phones.

The significance of Pipette is that it measures the conditions that determine whether local AI is usable on a particular device. A model may perform well on a conventional capability benchmark yet be too slow, memory-intensive or thermally demanding for a phone. Latency affects interactive use, throughput affects how quickly text can be generated, and memory use determines whether a model can run alongside the operating system and other applications. By exposing these measures, a tool such as Pipette can give developers more practical information when selecting models for offline or privacy-sensitive applications.

BigGo reports that conventional benchmarks are often designed around large, data-center-class models. According to the report, smaller models running on smartphones sometimes fail to complete those tests and receive scores of zero, while smartphone inference software is less mature and differs from data-center infrastructure. A benchmark tailored to small, quantized models may therefore provide a more useful picture of device-level tradeoffs. That does not make the results directly comparable with scores from larger systems. It means the evaluation is answering a narrower question: how capable and responsive a model is under specified local hardware and software conditions.

The reported results also illustrate why speed and intelligence cannot be treated as a single metric. BigGo says Nanbeige4.2-3B scored similarly to LFM2.5-2.6B but tended to require more time to produce 256 tokens. It also reports relatively low token efficiency for Qwen3.5 9B Reasoning and Qwen3.5 4B Reasoning in one analysis. These findings could matter for developers choosing between a more capable model and a faster one, but they should be treated as results from the reported test setup rather than general judgments about the models. No independent evaluation, user study or production deployment evidence is provided.

Interactive Mechanism

Cơ chế tương tác: Nó thực sự hoạt động như thế nào

Khám phá công nghệ cơ bản đằng sau sự phát triển này một cách tương tác.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Kiểm tra khái niệm tương tác+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Xem gì tiếp theo

The main limitation is comparability. BigGo reports that iOS currently supports GPU acceleration while Android is CPU-only, and Liquid AI warns against cross-device comparisons. Future updates are expected to add models and comparisons based on memory usage, but the source does not independently establish the app’s broader reliability, adoption or long-term roadmap.

The first issue to watch is platform parity. BigGo reports that Pipette’s iOS version supports GPU-accelerated computation while the Android version is currently CPU-only. The report also identifies differences in flash-attention support, thread counts, accelerators and execution environments. Those differences can dominate performance results, so a score from an iPhone 17 Pro should not be read as a ranking of Android devices or as a direct comparison between operating systems. Liquid AI’s warning against cross-device comparisons is an important qualification.

The second issue is methodological transparency and reproducibility. Pipette can export results as CSV, according to BigGo, and its source code is publicly available. That could allow developers to inspect test configurations and repeat measurements. However, the source does not establish whether all model files, runtime versions, thermal conditions, compiler settings and benchmark implementations are identical across tests. It also does not report independent replication by researchers or users. Those details will determine how much confidence should be placed in differences of only a few tokens per second or small score changes.

Finally, the reported roadmap could broaden the tool’s usefulness if implemented as described. BigGo says Artificial Analysis plans to add models over time and compare models using memory usage rather than relying on one precision. That would help users evaluate tradeoffs among model size, quality, speed and resource consumption. Important unknowns remain: the source does not establish how widely Pipette is being used, whether the Android implementation will gain GPU support, how results are moderated or audited, or whether the benchmark predicts performance in real applications. Until those questions are answered, Pipette is best understood as a practical measurement tool for specific configurations, not a definitive leaderboard for smartphone AI.

Hướng dẫn và câu hỏi liên quan

Giải thích về mô hình AIĐào tạo AIMáy biến ápKiểm tra những gì bạn biết — thử một bài kiểm tra AI miễn phíTra cứu một thuật ngữ AI trong bảng thuật ngữ của chúng tôiTheo dõi trình theo dõi phát hành mô hình AI

Cập nhật và sửa chữa

Câu chuyện kinh điển này được cập nhật tại chỗ khi sự kiện đang phát triển có thay đổi cơ bản. URL và ngày xuất bản ban đầu của nó không bao giờ thay đổi.

  • BigGo Finance materially advances the existing Pipette report with a hands-on iPhone 17 Pro workflow, a reported 48.3726-token-per-second Gemma 4 E2B IT result, an 18-minute test duration, details on the app’s four measurement modes and data-sharing option, and additional Artificial Analysis benchmark findings. The measurements and rankings are attributed to BigGo and are not independently confirmed here.
Xem nhật ký chỉnh sửa công khai
Tìm thấy điều này hữu ích?