hraness
Theme
Appearance

saved

SemanticFinder: frontend-only live semantic search

by do-meGitHub

Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.

gist

SemanticFinder is a browser-only semantic search and chat-with-your-documents app built on transformers.js, with Wasm and WebGPU backends. It embeds text segments and a query client-side, ranks them by cosine similarity, and highlights matches live without sending the document to a server. A catalogue of pre-indexed classics (Bible, Les Misérables, IPCC, and others) returns results in about two seconds; a Chrome extension and Hugging Face export path extend the same privacy-first pipeline.

ideas

  • Embeddings never leave the browser. transformers.js runs feature-extraction models locally; Wasm covers the model set, WebGPU is the fast path.
  • Segment size trades recall for speed. Longer chunks finish faster; only changing segment length forces a full re-embed because embeddings are cached per segment.
  • Pre-indexed catalogues beat cold starts. Shared JSON.gz indices for large books let subsequent queries finish in about two seconds after the model loads.
  • Hybrid search and a threshold gate results. Semantic plus full-text ranking, with a similarity cutoff deciding how many highlights appear.
  • For libraries, use a real vector DB. Moby Dick-scale one-shot indexing works in-browser; sustained multi-book libraries need an external vector store.

quotes

“Calculates the embeddings and cosine similarity client-side without server-side inferencing”

do-me, stating the client-side embedding thesis.

“Data privacy-friendly - your input text data is not sent to a server, it stays in your browser!”

do-me, naming the privacy default.

“Following queries take only ~2 seconds!”

do-me, reporting post-index query latency on Moby Dick.

“Only if the user changes the segment length, the embeddings must be recalculated.”

do-me, explaining the cache invalidation rule.