返回新聞
創新AI Understanding 簡報

MarkTechPost 報告 Keenable AI 開源 NEEDLE,這是搜尋 API 的即時基準

MarkTechPost 報導稱,Keenable AI 發布了 NEEDLE,這是一個開源基準測試,每小時或每天刷新一次搜尋查詢,以測試人工智慧搜尋系統是否可以在不依賴記憶答案集的情況下檢索當前和困難的資訊。

5 min readRead the linked source
Source-provided image accompanying MarkTechPost reports Keenable AI open-sources NEEDLE, a live benchmark for search APIs
來源參考來源記錄
出版商
marktechpost.com
來源連結
marktechpost.comhttps://www.marktechpost.com/2026/08/31/keenable-ai-open-sources-needle-a-live-search-benchmark-that-rebuilds-its-query-set-every-hour/
來源類型
連結來源-主要來源狀態尚未確定。
背景60 秒內了解這一點

從這裡開始

關鍵術語

基準測試
用於測量和比較模型性能的標準化測試或資料集。
API(應用程式介面)
一種軟體系統向另一個系統發送請求並接收回應的結構化方式。
大語言模型(LLM)
在海量文本語料庫上訓練來產生和分析文本的語言模型。
測試一下自己AI 代理測驗

發生了什麼事

MarkTechPost reports that Keenable AI released NEEDLE, an open-source for evaluating web-search APIs used by AI agents. Its query sets are regenerated from fresh public sources rather than fixed in advance, with news queries rebuilt hourly and finance, scholar, legal, and deep-tail queries rebuilt daily.

MarkTechPost reports that Keenable AI has open-sourced NEEDLE, whose name refers to News, Everyday, Expert, Deep-tail, and Legal Evaluation. The is intended for search systems used by AI agents, where a model may issue repeated queries and rely on titles, snippets, and rankings before deciding whether to fetch a page. The central design is to avoid a permanently fixed question set. According to the report, news queries are regenerated every hour from roughly 124 curated RSS feeds and Google Trends. Finance, scholar, legal, and rare-entity queries are regenerated daily from sources including SEC XBRL, Wikidata, GLEIF, arXiv, Europe PMC, CourtListener, eCFR, and public agent-trajectory releases.

MarkTechPost reports that an LLM converts source items into queries, while a machine check is used to detect whether the query itself gives away the answer. The covers five different task types. News and deep-tail searches are judged for ranking quality; finance tests whether a requested fact appears within the top five snippets; and scholar and legal searches are treated as known-item tasks scored by identifier matching. The article says the open-source harness is a Python command-line tool with generate and run subcommands, can run on a laptop or in continuous integration, and requires an OpenRouter key for judging plus an API key for each search engine tested.

For each query, MarkTechPost says NEEDLE sends the same query text to 15 search APIs under one protocol. Calls are made one at a time, allowing latency percentiles to be compared without concurrent load. The evaluation uses each engine’s own ranking, titles, and snippets; pages are not fetched, and results are not re-ranked. Evidence is clipped to 2,000 characters, and the judge is not shown the engine’s name. According to the report, public GitHub Actions runs and per-run artifacts on a Hugging Face dataset are part of the project’s reproducibility approach, though those resources are not independently verified in the supplied material.

來源詳情: marktechpost.com ↗

為什麼這很重要

The is designed to address a weakness in static search evaluations: an AI system may memorize published questions and answers, or retrieve a public answer key instead of actually searching. NEEDLE also compares each engine with an empirical ceiling based on the combined results returned by all tested providers.

The addresses a practical problem in measuring AI search. Static evaluations can become less informative when models have memorized the questions, when answers are embedded in training data, or when an agent can access a publicly available answer file. A continuously refreshed query set makes those shortcuts harder, at least in principle. MarkTechPost compares the approach with other live evaluations such as LiveBench, LiveCodeBench, and SWE-bench-Live, but the supplied article does not independently establish how NEEDLE’s methodology compares with those projects.

NEEDLE’s pooled-oracle measure is potentially useful because a low score can have different causes. MarkTechPost reports that the combines all engines’ returned results for a query, orders that combined pool by relevance, and uses it to create an empirical ceiling called “ultimate.” A large gap between an individual engine and that ceiling suggests that relevant material was retrieved by at least one provider but was not surfaced or ranked well by that engine. A weak ceiling suggests that the tested providers collectively failed to retrieve strong evidence. This is a diagnostic distinction rather than proof of overall search quality.

The reported results indicate that performance varies sharply by task. For the seven-day window ending August 28, MarkTechPost reports finance scores of 0.910 for Exa, 0.872 for Keenable, 0.871 for Perplexity, and 0.847 for Google, against an ultimate score of 0.965. Scholar scores ranged from 0.774 for Keenable to 0.310 for Tavily, against 0.869. Deep-tail was reported as the hardest category: Exa reached 0.557 of ultimate, Keenable 0.470, and Bing 0.199. These figures are claims from the outlet, not independently confirmed measurements.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
互動式概念檢查+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

接下來看什麼

The provided source does not independently confirm the repository, dashboard, dataset artifacts, or reported scores. Future runs should show whether the ’s live-source process remains reproducible, whether its pooled oracle meaningfully distinguishes retrieval failures from ranking failures, and how results change as search providers update their indexes and APIs.

The first issue to watch is reproducibility. MarkTechPost reports that NEEDLE’s code is released under the MIT license and that query streams can be recreated, but the supplied source does not include an independent audit of the code, source-ingestion schedule, leak checks, judge prompts, or published artifacts. Researchers and search providers will need to establish whether a fresh run produces the same query construction and whether changes in public feeds or registries are recorded clearly enough for later review.

The second issue is whether the “ultimate” ceiling is interpreted cautiously. It measures the best result found by the participating engines, not the complete universe of relevant web pages. It can therefore reveal shared misses only within the tested providers and source window. MarkTechPost also reports that NEEDLE’s operator, Keenable, competes in the . That creates a conflict that does not invalidate the method by itself, but makes transparent code, complete logs, blind judging, and independent reruns especially important.

Latency will also matter for agent developers. MarkTechPost reports Keenable-realtime at 193 milliseconds median and 284 milliseconds at the 95th percentile, compared with Exa at 1,876 and 2,955 milliseconds and Bing at 2,767 and 9,381 milliseconds for the same seven-day window. The article says failed calls are excluded from latency samples, so future evaluations should examine failure rates, rate limits, coverage changes, and the trade-off between speed and retrieval quality alongside these percentiles.

相關指引和測驗

人工智慧代理人工智慧模型解釋人工智慧培訓測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注 AI 模型發布追蹤器
覺得有用嗎?