Hraness
Theme
Appearance
← sponge

market · architecture · evidence

Sponge is a knowledge harness

Research reports are fast to produce with AI, but still costly to trust, reuse, and update. Sponge lets agents autonomously absorb their claims, evidence, and context into structured knowledge while people steer, review, and control release.

01 · market baseline

A cited research report now takes minutes

AI makes cited reports fast to produce. Sponge addresses the harder work of checking, reusing, and updating their contents.

Perplexity

~4–5 min
control
automatic path
sources
web · files · premium
artifact
PDF · DOCX · Page

ChatGPT

no estimate
control
editable plan + steering
sources
web · files · apps / MCP
artifact
Markdown · Word · PDF

Gemini

~5–10 min
control
editable plan
sources
Search · Drive · NotebookLM
artifact
Canvas · Docs · media

Claude

~1–3 min
control
parallel research agents
sources
web · Workspace · MCP
artifact
cited chat report

The 2026 baseline

Perplexity, ChatGPT, Gemini, and Claude all search public and private sources and produce cited reports. ChatGPT and Gemini let users edit the research plan; Claude dispatches parallel subagents; Perplexity chooses the path and models. Outputs include chat, Canvas, Markdown, Word, PDF, Docs, Pages, audio, and visuals.

Perplexity estimates four to five minutes, Gemini five to ten, and Anthropic one to three for Claude Research. OpenAI publishes no completion-time estimate. Quotas and capabilities vary by plan. See the documentation for Perplexity Research, ChatGPT deep research, Gemini Deep Research, and Claude Research.

Sponge starts where the report ends. Its job is to absorb the underlying claims, evidence, and context into durable records that people can steer, review, correct, and reuse.

02 · measured gap

DRACO: presentation criteria 77–94%; factual criteria 49–66%

Factual and citation quality still trail presentation. Experts must decide which claims can enter durable knowledge.

Perplexity DR (Opus 4.6)
66.4%
75.1%
93.8%
OpenAI DR (o3)
49.0%
60.4%
77.0%
Gemini DR
51.0%
64.4%
92.1%
Claude Opus 4.6*
54.4%
67.2%
83.1%

* standard Claude Opus 4.6 with web and code; Anthropic had no Research API

The audit burden remains

Perplexity’s February 2026 DRACO benchmark evaluates 100 difficult research tasks. Across the four API configurations shown here, presentation criteria pass 77.0–93.8% of the time while factual criteria pass 49.0–66.4%. The gap ranges from 27.4 to 41.1 percentage points. Citation-quality pass rates fall in between.

DRACO has important limits: Perplexity authored it; the tasks are single-turn, English, and text-only; LLM judges grade the outputs; and the Claude row uses ordinary Opus with web and code because Anthropic exposes no Research API. ResearchRubrics, accepted at ICLR 2026, evaluated 101 prompts against 2,593 expert-written criteria. The best system reached 61.5% compliance under binary grading. Implicit reasoning and synthesis accounted for roughly 45–50% of failures.

A report produced in minutes can require hours of expert review. Reviewers break conclusions into claims, check whether citations support them, search for missing counterevidence, and assess each source’s authority. DeepFact found that PhD specialists labeled hidden known-answer claims at 60.8% accuracy unaided; after three evidence-backed audit and revision rounds, accuracy reached 90.9%. Sponge stores those claim-level checks as reusable review state.

03 · specialist systems

Research systems expose domain objects and review state

Specialist systems make domain objects and review state explicit. A knowledge harness must preserve both across sources.

Elicit

research agent + reviews

find → screen → extract → analyze

projects · sentence citations · artifacts

Covidence

systematic review

dual screen → resolve → extract

exclusions · conflicts · risk of bias

Benchling

wet-lab R&D

register → run → trace samples

entities · inventory · assays

Galaxy

computational biology

inputs → tools → workflow history

versions · provenance · rerun

Structure follows the job

Elicit Research Agent spans literature, clinical trials, patents, biology and regulatory databases, the web, and uploaded internal data. Projects retain sources, instructions, context, and outputs. Sentence-level citations support reports, tables, figures, calculations, documents, and slide decks. Elicit’s systematic-review workflow and Covidence record the protocol, search, deduplication, dual screening, exclusion reasons, extraction, risk-of-bias judgments, conflict resolution, and export.

Elicit’s May 6 vendor evaluation began with 994 unique open-access Cochrane reviews and applied a separate filter at each stage. It reports 95.0% included-study search recall, 96.9% abstract-screening sensitivity, 99.5% full-text recall, and 95.6% extraction accuracy. The evaluation write-up publishes the denominators and filters.

Benchling provides a biology-aware registry, ELN, inventory, result schemas, and workflows. Administrators must define schemas, naming rules, and registration constraints before a team can create governed entities. That setup costs time but enables deep biotech coverage. Galaxy captures inputs, tools, versions, histories, and workflows for rerunnable biomedical analysis.

Each system records different objects: Elicit and Covidence center papers and review decisions; Benchling centers biomolecules, samples, and experiments; Galaxy centers datasets, jobs, and histories. Their lineage crosses products through exports, APIs, or custom connectors. scite covers another slice by classifying citation statements as supporting, contrasting, or mentioning. Sponge aims to preserve these objects and review decisions as knowledge moves across sources and products.

04 · science platforms

Scientific AI now runs tools, code, databases, and lab loops

Science platforms preserve provenance inside their own products. Sponge would carry stable claim, evidence, and review IDs across them. Its first connector remains unbuilt.

OpenAI

  • Prismmanuscripts
    free
  • GPT-Rosalindlife sciences
    trusted access
  • 2 Life Sciences pluginsresearch + NGS
    all Codex users
  • 10,080 reactionsclosed-loop chemistry
    lab demonstration

Google

  • Literature InsightsNotebookLM
    experimental
  • Hypothesis GenerationCo-Scientist
    experimental
  • Computational DiscoveryERA + AlphaEvolve
    experimental
  • AlphaFoldstructure prediction
    available

Anthropic

  • Claude Scienceanalysis workbench
    public beta
  • 60+ databasesdomain tools
    public beta
  • Local · HPC · ModalPython / R compute
    public beta
  • Artifact + reviewercode · env · checks
    public beta

Current science stacks

OpenAI covers several stages. Prism is a free LaTeX-native writing workspace. The June GPT-Rosalind update reports 27.5% versus GPT-5.5’s 25.1% on OpenAI’s MedChemBench, 21.6% versus 20.4% on GeneBench, and 63.2% versus 55.8% on LabWorkBench. Two Life Sciences plugins, Research and NGS Analysis, are available to all Codex users and preserve artifacts and provenance in the workspace. In a separate Molecule.one experiment, models proposed work, an automated laboratory ran 10,080 reactions, and human chemists reproduced representative results.

Google’s May 2026 Gemini for Science suite separates Literature Insights, Hypothesis Generation, and Computational Discovery. Co-Scientist generates, debates, and ranks hypotheses. ERA and AlphaEvolve generate executable candidates and score them against an external metric. Access is gradual and experimental. AlphaFold Server is a separate mature service for biomolecular inputs, predicted complexes, and confidence measures.

Claude Science is a public beta for macOS and Linux on Pro, Max, Team, and Enterprise. It runs persistent Python and R kernels locally, over SSH on HPC, or through Modal. It connects to more than 60 scientific databases, binds figures, tables, and notebooks to their exact code, environment, and conversation, and flags citation, number, and figure/code mismatches with a background reviewer.

These systems preserve substantial provenance inside their own workbenches. Sponge would carry the same claim, evidence, and review IDs across products and domain packs. Its first connector remains unbuilt.

05 · open-source kernel

Oh v1 separates a claim from its attributed stance

Oh defines the harness’s versioned records and dependencies. Integrating @hraness/oh remains work ahead.

open full-size diagramopen full-size diagram

The open-source kernel

Oh is the open-source kernel Sponge is converging on. As of August 27, 2026, Sponge contains a richer implementation of the same seven-concept model and has not yet integrated @hraness/oh.

In Oh’s current ontology v1 specification, a biomarker is an Entity; “biomarker X predicts response” is an immutable Statement; and a study’s authors express a supporting Assertion in a trial-cohort Context. An Evidence link points to the exact passage or table cell and records how it bears on the assertion. A reproducible review section can be stored as a view record, which corresponds to the Projection concept.

The MIT-licensed hraness/oh v0.1.1 release, published August 27, 2026, stores content-addressed records and an append-only operation chain in SQLite. It supports generation-checked writes, replay verification, keyword search, optional local semantic search, and bounded fast-forward sync. Its CLI, SDK, and Agent Skill share one contract. Verification covers canonical bytes, dependencies, and history. Scientific validity remains a review decision. Extraction, citation ingestion, collaboration UI, domain packs, and executable SHACL validation remain application work.

Typed ontology parsers enforce assertion and evidence laws when an application invokes them. The generic CLI validates canonical envelopes and dependencies but does not automatically apply a domain codec.

06 · domain specialization

A biomedical pack would add expert fields and validation

Domain packs are the planned bridge from generic records to source-specific fields, vocabularies, evidence rules, and executable validation.

open full-size diagramopen full-size diagram

Specialization is the next build

Oh v1 defines schema-revision kinds for concepts, predicates, mappings, shapes, units, and vocabularies, plus graph kinds for shape and type-membership. Applications supply semantic codecs; Oh’s generic store does not execute those records. A domain pack could add that layer. An assay record could require specimen, organism or model, protocol, instrument, measured variable, unit, and source locator. A clinical effect estimate could require cohort, comparator, endpoint, time horizon, value, uncertainty interval, and the exact table cell or figure region.

A biomedical pack could reuse OBI for investigations and assays, GO and ECO for functional annotations and evidence codes, DICOM identifiers for imaging, SNOMED CT for clinical findings, and LOINC for laboratory tests. OBO Foundry principles set expectations for persistent identifiers, versioned releases, defined scope, and relation reuse. SHACL defines machine-checkable graph constraints.

Other packs would model different objects: definitions, lemmas, proof dependencies, and counterexamples for mathematics; jurisdiction, instrument, effective date, amendment, and supersession for policy. Pack-defined mappings must specify cross-domain identity and derivation semantics. Today’s Oh release includes the extension records. The first domain pack and executable constraint evaluator remain unbuilt, as the diagram shows.

07 · current system

Sponge adds a workbench and review gates

Today, Sponge lets agents edit bounded document blocks and propose graph changes while named people retain acceptance and publication authority.

open full-size diagramopen full-size diagram

Human steering and review

The current Sponge workbench implements human steering and review. It gives document blocks stable IDs for their text, citations, links, outline position, and history. An agent can commit a fenced private edit and activate a separate authority-free graph proposal. The authenticated owner reviews each change for a stated purpose. Public use requires public-purpose review. Finalization creates an immutable edition; a later publish action creates its public locator.

The API, CLI, and installable Agent Skill expose the same objects. Credentials remain scoped and revocable. An agent can search, draft, and submit commands without gaining release authority. The public Sponge docs and developer guide describe the workbench and programmatic interfaces.

Oh is the open-source target for canonical records, operation history, local authority, verification, search, and sync. Sponge currently contains a richer implementation of the shared model. Replacing that internal kernel with @hraness/oh is the next engineering step. The workbench, agent coordination, and review lifecycle remain Sponge application code.

08 · proposed pilot

Pilot: one oncology biomarker, one indication, one evidence packet

The proposed pilot would test the harness on one public evidence packet with exact source coordinates and measured human review.

open full-size diagramopen full-size diagram

A narrow first wedge

A useful first pilot is narrower than a general clinical platform: one oncology biomarker, one indication, published papers, supplementary tables and figures, a public ClinicalTrials.gov record, and any correction or retraction notices. Its versioned evidence packet records the trial, cohort, intervention, comparator, endpoint, effect estimate, uncertainty, adverse events, and an exact source coordinate for every field.

The first study uses public aggregate evidence. Patient-linked DICOM, whole-slide, and molecular integration belong in a later governed phase. ClinicalTrials.gov support is metadata-only today. Domain-linkage schemas, evidence rubrics, contradiction ledgers, and clinical governance remain unfinished. The pilot stays researcher-facing and measures review time, extraction corrections, and unresolved disagreements.