Quay lại Tin tức
Đổi mớiAI Understanding tóm tắt

MarkTechPost báo cáo nguồn mở Keenable AI NEEDLE, điểm chuẩn trực tiếp cho API tìm kiếm

MarkTechPost báo cáo rằng Keenable AI đã phát hành NEEDLE, một điểm chuẩn nguồn mở giúp làm mới các truy vấn tìm kiếm hàng giờ hoặc hàng ngày để kiểm tra xem hệ thống tìm kiếm AI có thể truy xuất thông tin hiện tại và khó mà không cần dựa vào các bộ câu trả lời được ghi nhớ hay không.

5 min readRead the linked source
Source-provided image accompanying MarkTechPost reports Keenable AI open-sources NEEDLE, a live benchmark for search APIs
Nguồn tham khảoNguồn đã ghi
Nhà xuất bản
marktechpost.com
Liên kết nguồn
marktechpost.comhttps://www.marktechpost.com/2026/08/31/keenable-ai-open-sources-needle-a-live-search-benchmark-that-rebuilds-its-query-set-every-hour/
Loại nguồn
Nguồn được liên kết - trạng thái nguồn chính chưa được thiết lập.
Bối cảnhHiểu điều này trong 60 giây

Bắt đầu ở đây

Thuật ngữ chính

Điểm chuẩn
Một bài kiểm tra hoặc tập dữ liệu được tiêu chuẩn hóa dùng để đo lường và so sánh hiệu suất của mô hình.
API (Giao diện lập trình ứng dụng)
Một cách có cấu trúc để một hệ thống phần mềm gửi yêu cầu và nhận phản hồi từ hệ thống khác.
Mô hình ngôn ngữ lớn (LLM)
Một mô hình ngôn ngữ được đào tạo trên kho văn bản lớn để tạo và phân tích văn bản.
Tự kiểm traCâu đố về đại lý AI

Chuyện gì đã xảy ra

MarkTechPost reports that Keenable AI released NEEDLE, an open-source for evaluating web-search APIs used by AI agents. Its query sets are regenerated from fresh public sources rather than fixed in advance, with news queries rebuilt hourly and finance, scholar, legal, and deep-tail queries rebuilt daily.

MarkTechPost reports that Keenable AI has open-sourced NEEDLE, whose name refers to News, Everyday, Expert, Deep-tail, and Legal Evaluation. The is intended for search systems used by AI agents, where a model may issue repeated queries and rely on titles, snippets, and rankings before deciding whether to fetch a page. The central design is to avoid a permanently fixed question set. According to the report, news queries are regenerated every hour from roughly 124 curated RSS feeds and Google Trends. Finance, scholar, legal, and rare-entity queries are regenerated daily from sources including SEC XBRL, Wikidata, GLEIF, arXiv, Europe PMC, CourtListener, eCFR, and public agent-trajectory releases.

MarkTechPost reports that an LLM converts source items into queries, while a machine check is used to detect whether the query itself gives away the answer. The covers five different task types. News and deep-tail searches are judged for ranking quality; finance tests whether a requested fact appears within the top five snippets; and scholar and legal searches are treated as known-item tasks scored by identifier matching. The article says the open-source harness is a Python command-line tool with generate and run subcommands, can run on a laptop or in continuous integration, and requires an OpenRouter key for judging plus an API key for each search engine tested.

For each query, MarkTechPost says NEEDLE sends the same query text to 15 search APIs under one protocol. Calls are made one at a time, allowing latency percentiles to be compared without concurrent load. The evaluation uses each engine’s own ranking, titles, and snippets; pages are not fetched, and results are not re-ranked. Evidence is clipped to 2,000 characters, and the judge is not shown the engine’s name. According to the report, public GitHub Actions runs and per-run artifacts on a Hugging Face dataset are part of the project’s reproducibility approach, though those resources are not independently verified in the supplied material.

Chi tiết nguồn: marktechpost.com ↗

Tại sao nó quan trọng

The is designed to address a weakness in static search evaluations: an AI system may memorize published questions and answers, or retrieve a public answer key instead of actually searching. NEEDLE also compares each engine with an empirical ceiling based on the combined results returned by all tested providers.

The addresses a practical problem in measuring AI search. Static evaluations can become less informative when models have memorized the questions, when answers are embedded in training data, or when an agent can access a publicly available answer file. A continuously refreshed query set makes those shortcuts harder, at least in principle. MarkTechPost compares the approach with other live evaluations such as LiveBench, LiveCodeBench, and SWE-bench-Live, but the supplied article does not independently establish how NEEDLE’s methodology compares with those projects.

NEEDLE’s pooled-oracle measure is potentially useful because a low score can have different causes. MarkTechPost reports that the combines all engines’ returned results for a query, orders that combined pool by relevance, and uses it to create an empirical ceiling called “ultimate.” A large gap between an individual engine and that ceiling suggests that relevant material was retrieved by at least one provider but was not surfaced or ranked well by that engine. A weak ceiling suggests that the tested providers collectively failed to retrieve strong evidence. This is a diagnostic distinction rather than proof of overall search quality.

The reported results indicate that performance varies sharply by task. For the seven-day window ending August 28, MarkTechPost reports finance scores of 0.910 for Exa, 0.872 for Keenable, 0.871 for Perplexity, and 0.847 for Google, against an ultimate score of 0.965. Scholar scores ranged from 0.774 for Keenable to 0.310 for Tavily, against 0.869. Deep-tail was reported as the hardest category: Exa reached 0.557 of ultimate, Keenable 0.470, and Bing 0.199. These figures are claims from the outlet, not independently confirmed measurements.

Interactive Mechanism

Cơ chế tương tác: Nó thực sự hoạt động như thế nào

Khám phá công nghệ cơ bản đằng sau sự phát triển này một cách tương tác.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
Kiểm tra khái niệm tương tác+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

Xem gì tiếp theo

The provided source does not independently confirm the repository, dashboard, dataset artifacts, or reported scores. Future runs should show whether the ’s live-source process remains reproducible, whether its pooled oracle meaningfully distinguishes retrieval failures from ranking failures, and how results change as search providers update their indexes and APIs.

The first issue to watch is reproducibility. MarkTechPost reports that NEEDLE’s code is released under the MIT license and that query streams can be recreated, but the supplied source does not include an independent audit of the code, source-ingestion schedule, leak checks, judge prompts, or published artifacts. Researchers and search providers will need to establish whether a fresh run produces the same query construction and whether changes in public feeds or registries are recorded clearly enough for later review.

The second issue is whether the “ultimate” ceiling is interpreted cautiously. It measures the best result found by the participating engines, not the complete universe of relevant web pages. It can therefore reveal shared misses only within the tested providers and source window. MarkTechPost also reports that NEEDLE’s operator, Keenable, competes in the . That creates a conflict that does not invalidate the method by itself, but makes transparent code, complete logs, blind judging, and independent reruns especially important.

Latency will also matter for agent developers. MarkTechPost reports Keenable-realtime at 193 milliseconds median and 284 milliseconds at the 95th percentile, compared with Exa at 1,876 and 2,955 milliseconds and Bing at 2,767 and 9,381 milliseconds for the same seven-day window. The article says failed calls are excluded from latency samples, so future evaluations should examine failure rates, rate limits, coverage changes, and the trade-off between speed and retrieval quality alongside these percentiles.

Hướng dẫn và câu hỏi liên quan

Đại lý AIGiải thích về mô hình AIĐào tạo AIKiểm tra những gì bạn biết — thử một bài kiểm tra AI miễn phíTra cứu một thuật ngữ AI trong bảng thuật ngữ của chúng tôiTheo dõi trình theo dõi phát hành mô hình AI
Tìm thấy điều này hữu ích?