뉴스로 돌아가기
제품AI Understanding 브리핑

Liquid AI의 Pipette 앱은 스마트폰의 기기 내 AI 성능을 측정한다고 BigGo가 보고했습니다.

BigGo Finance는 Liquid AI가 스마트폰 AI 대기 시간, 처리량 및 메모리 사용을 측정하기 위한 무료 iOS 및 Android 앱인 Pipette를 출시했다고 보고했습니다. 이 보고서에는 iPhone 17 Pro 테스트가 기록되어 있으며 인공 분석을 통해 개발된 스마트폰 중심 벤치마크가 자세히 설명되어 있습니다.

5 min readRead the linked source
Source-provided image accompanying Liquid AI’s Pipette app measures on-device AI performance on smartphones, BigGo reports
소스 참조녹음된 소스
출판사
finance.biggo.com
소스 링크
finance.biggo.comhttps://finance.biggo.com/news/d74ec349-df7a-4f48-a861-4811784c123c
소스 유형
연결된 소스 — 기본 소스 상태가 설정되지 않았습니다.
또한 인용됨

마지막으로 수정된 스토리

맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

온디바이스 AI
AI 추론은 원격 클라우드 서비스가 아닌 사용자 하드웨어에서 로컬로 수행됩니다.
메모리(에이전트 메모리)
AI 에이전트는 연속성을 향상하기 위해 여러 단계 또는 세션에서 사용하는 저장된 컨텍스트입니다.
양자화
모델 가중치를 8비트 또는 4비트와 같은 낮은 정밀도 형식으로 변환합니다.
자신을 테스트해 보세요AI 모델 설명 퀴즈

출간 이후 달라진 점

  1. 처음 출판됨
  2. BigGo Finance materially advances the existing Pipette report with a hands-on iPhone 17 Pro workflow, a reported 48.3726-token-per-second Gemma 4 E2B IT result, an 18-minute test duration, details on the app’s four measurement modes and data-sharing option, and additional Artificial Analysis benchmark findings. The measurements and rankings are attributed to BigGo and are not independently confirmed here.

무슨 일이 일어났나요?

Liquid AI released Pipette, a free smartphone app that lets users run quantized AI models locally and measure execution performance. BigGo Finance tested the iOS version on an iPhone 17 Pro and recorded 48.3726 tokens per second when running a 4-bit quantized Gemma 4 E2B IT model with MLX. BigGo also reported benchmark results from an Artificial Analysis collaboration covering small models that fit within 8GB when quantized.

BigGo Finance reports that Liquid AI released Pipette for iOS and Android as a free benchmark app for measuring AI model execution on smartphones. Users can download quantized models, including models from the Gemma and Qwen families, and run local tests for end-to-end latency, prefill throughput, decode throughput and maximum memory usage. The app requires account registration at first launch, according to BigGo, and downloaded models are stored on the device for testing. The source code is publicly available on GitHub, but this report does not independently verify the contents of that repository or the app’s availability in every market.

BigGo’s hands-on account describes a workflow in which users create a job, select either llama.cpp or MLX, choose a model and select test types. The report says all four test types are selected by default. It also says users can choose whether results are sent to the server as public data. Testing includes cooldown periods intended to reduce device temperature, and BigGo says its test of a 4-bit quantized Gemma 4 E2B IT model with MLX on an iPhone 17 Pro took 18 minutes. The resulting decode-throughput figure was 48.3726 tokens per second. These observations are reported by BigGo and have not been independently confirmed here.

BigGo also describes a separate collaboration between Artificial Analysis and Liquid AI to evaluate small models on an iPhone 17 Pro. The reported system combines Artificial Analysis’s intelligence-performance tests with Liquid AI’s inference-performance testing. Five benchmarks were used: BFCL, IFBench, AA-Omniscience, GPQA Diamond and MATH-500. BigGo says the evaluation was limited to models fitting within 8GB at 4-bit and used a context-length cap of 16,000 tokens for average-score comparisons. Nanbeige4.2-3B and LFM2.5-2.6B reportedly achieved the highest average scores under those conditions, while LFM2.5-8B-A1B led a comparison capped at one minute of response time. The source provides no independently checked raw benchmark table or replication results.

소스 세부정보: finance.biggo.com ↗

왜 중요한가요?

Pipette makes practical constraints more visible by measuring latency, throughput and memory usage on the hardware where models actually run. That can help developers and users distinguish theoretical model quality from the speed and resource demands of local inference, although the results are platform-specific and not a universal ranking of models or phones.

The significance of Pipette is that it measures the conditions that determine whether local AI is usable on a particular device. A model may perform well on a conventional capability benchmark yet be too slow, memory-intensive or thermally demanding for a phone. Latency affects interactive use, throughput affects how quickly text can be generated, and memory use determines whether a model can run alongside the operating system and other applications. By exposing these measures, a tool such as Pipette can give developers more practical information when selecting models for offline or privacy-sensitive applications.

BigGo reports that conventional benchmarks are often designed around large, data-center-class models. According to the report, smaller models running on smartphones sometimes fail to complete those tests and receive scores of zero, while smartphone inference software is less mature and differs from data-center infrastructure. A benchmark tailored to small, quantized models may therefore provide a more useful picture of device-level tradeoffs. That does not make the results directly comparable with scores from larger systems. It means the evaluation is answering a narrower question: how capable and responsive a model is under specified local hardware and software conditions.

The reported results also illustrate why speed and intelligence cannot be treated as a single metric. BigGo says Nanbeige4.2-3B scored similarly to LFM2.5-2.6B but tended to require more time to produce 256 tokens. It also reports relatively low token efficiency for Qwen3.5 9B Reasoning and Qwen3.5 4B Reasoning in one analysis. These findings could matter for developers choosing between a more capable model and a faster one, but they should be treated as results from the reported test setup rather than general judgments about the models. No independent evaluation, user study or production deployment evidence is provided.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
대화형 개념 확인+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

다음에 무엇을 볼 것인가

The main limitation is comparability. BigGo reports that iOS currently supports GPU acceleration while Android is CPU-only, and Liquid AI warns against cross-device comparisons. Future updates are expected to add models and comparisons based on memory usage, but the source does not independently establish the app’s broader reliability, adoption or long-term roadmap.

The first issue to watch is platform parity. BigGo reports that Pipette’s iOS version supports GPU-accelerated computation while the Android version is currently CPU-only. The report also identifies differences in flash-attention support, thread counts, accelerators and execution environments. Those differences can dominate performance results, so a score from an iPhone 17 Pro should not be read as a ranking of Android devices or as a direct comparison between operating systems. Liquid AI’s warning against cross-device comparisons is an important qualification.

The second issue is methodological transparency and reproducibility. Pipette can export results as CSV, according to BigGo, and its source code is publicly available. That could allow developers to inspect test configurations and repeat measurements. However, the source does not establish whether all model files, runtime versions, thermal conditions, compiler settings and benchmark implementations are identical across tests. It also does not report independent replication by researchers or users. Those details will determine how much confidence should be placed in differences of only a few tokens per second or small score changes.

Finally, the reported roadmap could broaden the tool’s usefulness if implemented as described. BigGo says Artificial Analysis plans to add models over time and compare models using memory usage rather than relying on one precision. That would help users evaluate tradeoffs among model size, quality, speed and resource consumption. Important unknowns remain: the source does not establish how widely Pipette is being used, whether the Android implementation will gain GPU support, how results are moderated or audited, or whether the benchmark predicts performance in real applications. Until those questions are answered, Pipette is best understood as a practical measurement tool for specific configurations, not a definitive leaderboard for smartphone AI.

관련 가이드 및 퀴즈

AI 모델 설명AI 트레이닝트랜스포머알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.AI 모델 출시 추적기를 따르세요.

업데이트 및 수정

이 정식 스토리는 진행 중인 이벤트가 실질적으로 변경될 때 업데이트됩니다. URL과 원래 출판 날짜는 절대 변경되지 않습니다.

  • BigGo Finance materially advances the existing Pipette report with a hands-on iPhone 17 Pro workflow, a reported 48.3726-token-per-second Gemma 4 E2B IT result, an 18-minute test duration, details on the app’s four measurement modes and data-sharing option, and additional Artificial Analysis benchmark findings. The measurements and rankings are attributed to BigGo and are not independently confirmed here.
공개 수정 로그 보기
이것이 유용하다고 생각하시나요?