Teaching an Agent to Search by Ear, Without a Cloud API Call (3)
The last article covered what an agent sees once it's looking at a sound: the structured provenance record, the license fields it can trust, the story behind how the sound was made. None of that helps if the agent can't find the sound in the first place. That's what this article is about: how search actually works underneath search_sounds, and how it runs entirely on the same machine as the catalog, with no query ever leaving it.
Two ways search goes wrong
Keyword search is the obvious first move, and it fails in a specific, predictable way: it only finds what you asked for in the exact words you used. Search for "metallic" and a sound tagged only "clang" won't show up, even though it's probably what you wanted. That's a real gap for a catalog where the same idea gets described a dozen different ways across a dozen different sounds.
The obvious fix is semantic search, matching by meaning instead of exact words, and it has the opposite failure mode. Pushed too far on its own, it starts returning things that are vaguely in the neighborhood of the query without actually being right, because "close in meaning" and "correct" aren't the same thing. Neither approach alone is good enough. Ours runs both, every time, and combines the results.
The keyword half: SQLite, on purpose
For exact-term matching we use SQLite's FTS5 extension. That might sound like an odd choice next to Postgres with pgvector or a dedicated search engine like Meilisearch, and it's a deliberate one. Catalog sizes here run somewhere between a handful of sounds and a few tens of thousands, not the scale that would justify running and monitoring a separate database server. SQLite already does keyword search and vector search well within that range, in one file, with nothing else to deploy. We'd rather operate zero extra services than have a more impressive-sounding stack for a workload that doesn't need one.
The meaning half: embeddings, computed locally
For semantic matching, every sound's description gets converted into a vector, a list of numbers positioned so that similar meanings end up near each other in that space, using a small model called MiniLM (384 dimensions, run through a library called fastembed, no GPU required). That conversion happens once, at ingest time, and gets stored alongside the rest of the index.
The part worth pausing on: this runs entirely on the machine hosting the catalog. No description, and no search query, ever gets sent to OpenAI, Anthropic, or anyone else to be turned into a vector. That's not a privacy footnote we added later. It's one of the project's actual design goals, the idea that a catalog owner should be able to self-host the whole thing on infrastructure they control, with no per-query cost and no third party ever seeing what's being searched for.
Combining the two without hand-tuned weights
Once you have two ranked lists (one from keyword matching, one from vector similarity) you need a way to merge them that doesn't require guessing how much to trust each one. We use reciprocal rank fusion, RRF for short, and the idea is simpler than the name suggests: a result's score depends on where it ranked in each list, not on the raw numbers each method produced. A sound that finished 1st in the keyword search and 2nd in the vector search scores higher than one that only showed up in a single list at all, regardless of how confident either individual method claimed to be.
Here's what that looks like with real numbers. Say a query produces these two ranked lists:
Rank FTS5 branch Vector branch
1 glk-plate metal-clang
2 metal-clang glk-plate
3 rain-loop wind-loopglk-plate and metal-clang each show up in the top two of both lists, just swapped, so they end up tied for first after fusion. rain-loop and wind-loop each only appear in one list, in third place, so they both trail well behind. That's the property that makes RRF useful here: a sound both methods agree on, even moderately, consistently beats a sound only one method liked a lot. Chasing a single method's top pick would be easy to fool; RRF rewards agreement instead.
An honest caveat
Here's something we found while testing this against real queries, and it's worth stating plainly rather than glossing over. search_sounds has no minimum relevance score. Ask it for "everything recorded with a Schoeps microphone" against a catalog that contains zero microphone recordings at all, and it still comes back with ten results, ranked, looking reasonably confident about it.
That's not a bug. Every sound is a valid candidate unless a filter excludes it, and once RRF has a set of candidates to rank, it ranks them; there's no floor below which it just says "none of these are good enough." Scores are only ever meaningful relative to each other within one query, never as an absolute signal, and that's stated plainly in how the tool describes itself to an agent. But it does mean an agent, or a person, shouldn't treat appearing in a result list as proof of a specific claim. "Closest available" and "actually matching" aren't the same guarantee, and the honest fix isn't in the search code, it's in checking a specific claim against a sound's actual provenance record before repeating it.
What this sets up
Search only matters if there's a safe, well-defined way for an agent to actually call it, without also being able to do anything it shouldn't. That's the next article: the five tools the server exposes, why none of them can write anything, and what actually stops things from going wrong when you hand an AI agent direct access to a production catalog.