Powered by BM25 ranking, 120+ semantic synonyms, and multi-signal scoring — all running client-side with zero dependencies and zero network requests.
Introduction
Every developer knows the frustration: you’re working with a large codebase, and you need to find a specific piece of functionality. You know it exists somewhere — maybe in a utility file, maybe buried in a component — but plain text search just doesn’t cut it. You search for “auth” and get 200 results, but none of them are the login middleware you’re looking for.
Traditional code search tools treat your query as a simple string matching exercise. They have no understanding of what you’re actually looking for. Search for “password check” and you’ll miss the function called validateCredentials() — even though it does exactly what you need.
“What if your code search engine understood that ‘user login’, ‘sign in’, and ‘authenticate’ all mean the same thing — without sending your proprietary code to a cloud server?”
That’s exactly what we built. A semantic code search engine that runs entirely in your browser — ~600 lines of TypeScript, zero dependencies, 120+ synonym mappings, BM25 ranking, and multi-signal scoring. It indexes your code locally, understands the semantic meaning of your queries, and returns ranked, relevant results in under 200ms. All without a single network request.
The Problem
Ctrl+F is Not Enough
Plain text matching can’t understand synonyms, intent, or context. It returns every file containing the search terms, flooding you with irrelevant results.
Cloud Tools Leak Your Code
AI-powered code search tools send your proprietary source code to remote servers. For companies with strict compliance requirements, that’s a dealbreaker.
Language Servers are Heavy
LSP-based tools require installing language servers, configuring workspaces, and consuming significant RAM. They’re overkill for quick code search.
How It Works
The search engine follows a four-stage pipeline that transforms raw source code into ranked, relevant search results — all running entirely in the browser.
Parse
The code parser breaks source files into logical chunks — functions, classes, interfaces, methods, variables, and type definitions. Supports 15+ programming languages with automatic detection.
Tokenize
A smart tokenizer splits identifiers using camelCase, PascalCase, snake_case, and kebab-case conventions. It extracts meaningful tokens from function names, variable names, and docstrings while filtering noise.
Expand
The semantic expander maps natural language terms to programming concepts across 18 concept groups— auth, CRUD, API, UI, database, and more. Searching for “login” automatically expands to include “authenticate”, “signin”, “token”, “session”, and more.
Rank
A BM25-inspired ranking algorithm combined with multi-signal scoring delivers relevant results. Exact matches, identifier matches, semantic expansion hits, and docstring relevance all contribute to the final score.
Key Features
Zero Dependencies
No npm packages, no build tools, no external services. The entire engine is ~600 lines of pure TypeScript that runs in any modern browser. Drop it into a project and it just works.
Semantic Intelligence
120+ synonym mappings across 18 concept groups let you search using natural language. “How do users sign in?” finds authentication functions automatically.
Privacy First
Your code never leaves your browser. No network requests, no telemetry, no data collection. Perfect for proprietary codebases and compliance-sensitive environments.
Multi-Language
Automatic language detection for 15+ languages including TypeScript, Python, Rust, Go, Java, Ruby, and more. Upload entire projects and search across all files at once.
Smart Tokenizer
The tokenizer goes beyond simple whitespace splitting. It understands identifier conventions common in programming languages — splitting camelCase and PascalCase at word boundaries, decomposing snake_case and kebab-case, and filtering out common noise words.
BM25 Ranking Engine
We use a BM25-inspired scoring algorithm — the same family of algorithms that power full-text search engines like Elasticsearch and Lucene. IDF values are cached for performance, and the algorithm handles both exact and fuzzy matching.
Multi-Signal Scoring
Instead of relying on a single metric, the engine combines multiple signals to produce a comprehensive relevance score:
Performance Stats
~600 Lines of TypeScript
120+ Semantic Synonyms
15+ Languages Supported
<200ms Search Latency
Ready to Try It?
Launch the search engine and index your own code. No sign-up, no installation, no data leaving your browser.
Here is the code https://github.com/kalyanraju/semantic-search
0 Comments