hraness

saved

Grep achieves SOTA on major deep research benchmarks

by Grep AIXpublished

Hraness cites a source capture. The source author remains the source.

Grep AI @grepdotai

https://t.co/Xzo0oPrvMg automates high-stakes, repetitive work you can't afford to get wrong.

Grep achieves SOTA on major deep research benchmarks

Today we're sharing results from our evaluation of Grep across three major independent deep research benchmarks. Grep achieves state-of-the-art (SOTA) performance on all three, outperforming systems from Perplexity, Google, Anthropic, Nvidia, and OpenAI.

No other system has achieved SOTA on all of these benchmarks.

We are also introducing Grep Brain, a new architecture that transforms Grep from a research tool into a persistent system that learns your work, remembers across sessions, and produces real deliverables - slides, spreadsheets, reports, dashboards - not just research summaries.

All results, scoring code, and per-question data are published at github.com/Parcha-ai/benchmarks

Why this matters

Deep research is the foundation of serious work - the kind where being wrong has consequences and where you need to show your sources.

We built Grep out of what we learned running compliance infrastructure at Parcha, processing hundreds of thousands of requests for fintechs like Airwallex, Flutterwave, and IG.com at 99.7% accuracy. The same principles - source verification, structured reasoning, domain expertise - apply to any knowledge work where the stakes are real: compliance and AML, investor due diligence, underwriting, legal research, market intelligence, and OSINT.

Since launching as a limited research preview, Grep is now trusted by professionals at Amazon, Google, Citizens Bank, Orbital, Astreya, BitcoinSuisse, Craft.co, EffectGroup, Shopmonkey and dozens more companies to make high-stakes decisions.

The Results

DRACO

DRACO is Perplexity's benchmark, created with Harvard Business School. 100 open-ended research tasks across 10 domains, scored on factual accuracy, breadth, presentation, and citation quality.

Grep wins 9 of 10 domains. The largest margins are in the areas that require careful methodology: UX Design (+14.8pp), Needle in a Haystack (+12.4pp), Shopping/Product (+10.5pp). These are exactly the tasks where structured planning matters more than single-pass search.

Grep leads on all four rubric axes, with the biggest gap in Citation (+14.5pp) - a direct result of how the architecture tracks evidence through file-based context.

The gap between Grep and the same frontier model without deep research (Claude Opus 4.6 standalone) is 18.8 percentage points. Same model. Radically different results. That delta is architecture.

DeepSearchQA

Google DeepMind's benchmark. 896 multi-hop search questions requiring several steps of reasoning to arrive at a specific factual answer. Independently reproduced by Kaggle.

We achieved this with our medium-effort agent, not our highest. In Grep, effort doesn't just mean doing more work - it means being more focused. Excessive retrieval can actually hurt factual-correctness scores. The planning step calibrates effort to the task.

The floor matters more than the ceiling here. The lowest-performing category (Finance & Economics) still scores 79.5% on 132 questions. The spread from best to worst is only 13 points. Grep generalizes because planning adapts per question, rather than relying on a fixed retrieval pattern.

DeepResearch Bench

An independent benchmark from researchers in China from the University of Science and Technology of China, and Metastone Technology. 100 PhD-level questions (50 English, 50 Chinese), scored against expert-written reference articles using the RACE framework.

Grep leads in Instruction Following and Readability, and ranks top-3 across all dimensions against a field of 34 systems.

One detail stood out: Grep scored 56.42 on Chinese queries and 56.13 on English, despite zero optimization for Chinese. When the architecture is right, you don't need to tune for every language separately.

We ran three identical evaluations to measure scoring variance. Standard deviation: ~0.12. The results are stable.

Note: Grep's previous DeepResearch Bench submission is listed in second place with a score of 56.09. We have resubmitted our SOTA results and are awaiting official verification (https://huggingface.co/spaces/muset-ai/DeepResearch-Bench-Leaderboard).

How it works

Grep is one agent that reshapes itself across stages of research - and when the work demands it, splits into parallel copies of itself.

It is not a multi-agent system in the usual sense. There is no collection of specialized agents negotiating with each other. One agent morphs depending on the phase, loading different skills and tools, and spawns focused sub-agents when it needs to go wide. Think cell division, not a committee.

The planner does actual research before writing the plan. It gathers context, identifies gaps, and shapes the expert that will execute - selecting from 259 specialized skills and 90+ trusted data sources. The plan isn't a script. It's an opinionated view of how the work should proceed.

Execution runs as a two-layered loop. The inner loop spawns sub-agents, evaluates their findings, identifies gaps, and spawns more until coverage is sufficient. The outer loop reviews the finished report against quality criteria. If the report doesn't meet the bar, it goes back to research. The agent decides when it's done - not a timer, not a token budget.

What users say

"Our leadership asked if Gemini could do what Grep does. The team tried and it's not even close. Grep is far superior for our underwriting use case." – Crystal Anderson, Leading Risk Strategy & Operations, Shopmonkey

Shopmonkey's underwriting team uses Grep to make faster, more confident credit decisions across thousands of auto repair shops. In one case, Stripe was about to shut down a merchant for excessive disputes. Crystal ran the business through Grep. The report surfaced context that told a different story - the issues were linked to a closed business and a former employee. She told Stripe to stand down. The merchant was legitimate.

What's next

We are increasingly finding that what makes an expert effective is not more model capability but better infrastructure around it: specialized skills, a cloud-based computer to work on, and a brain that understands you and your work.

All benchmark results, scoring code, and per-question data is available here for review: github.com/Parcha-ai/benchmarks

Grep is available as a research preview to anyone who needs to use seriously deep research for making high-stakes decisions.

To try it Grep, Join the waitlist or schedule an onboarding call.

---

Written by Miguel Rios is co-founder and CTO of Grep. Previously Head of Platform Engineering at Brex and Head of Consumer Data Science at Twitter.

Cover image text:

GREP #1

Perplexity Benchmark

Google Benchmark

DeepResearch Bench

GREP

$1

Black-and-white circuit-board promotional graphic. Center trophy and box read GREP #1. Three windows labeled Perplexity Benchmark, Google Benchmark, and DeepResearch Bench each show GREP on the top podium with a checkmark; Perplexity and Google windows also show $1.