Photo by Stefan Steinbauer / Unsplash

Your Sound Library Is Invisible to AI (And That's a Problem) (1)

AI Aug 5, 2026

Ask an AI assistant to find you a metallic, low-pitched impact sound you can use commercially, and watch what happens. If it's connected to the web, it'll point you at a stock-sound marketplace. If it's connected to your DAW, it can load a track, route a bus, maybe even trigger playback. What it almost certainly can't do is look inside your own sound catalog, the one you built, curated, and licensed carefully, and just answer the question.

That catalog is invisible to it. Not on purpose. Nobody ever gave it a way in.

What agents can already do

AI agents have gotten genuinely good at operating on systems that expose an interface: a DAW's plugin API, a marketplace's public search endpoint. A growing set of community tools lets them control a DAW directly, creating tracks, placing clips, adjusting parameters. Other tools let them browse the open web, including the large generic sample marketplaces that happen to publish a search API. None of that is the problem.

The problem is everything an agent conspicuously can't reach: a specific publisher's own catalog, with its own metadata, its own licensing terms, its own story behind each sound. That catalog usually lives in a database only a website's search box can query, or worse, in a folder structure, a spreadsheet, a stack of PDFs, and whatever the person who built it happens to remember. An AI agent has no way in, not because the catalog is secret, but because nothing about it was ever built to be machine-callable.

Why this matters more for boutique catalogs than for stock platforms

For a generic stock-sound platform, "search by keyword" mostly gets you there; one impact sound is broadly interchangeable with another. For a boutique or independent catalog, that's usually not true. The actual product is often the story: how a sound was captured, what gear was used, how it was processed, exactly what the license does and doesn't permit. That story is the differentiator, and today it's almost never written down anywhere an agent can read it. If it exists at all, it's prose, aimed at a human.

An AI agent with access to a generic marketplace can find a sound. It can't tell you why this one is the right one, or whether you're actually allowed to use it, because nobody ever gave it that information in a form it can trust.

That gap gets more expensive the more people default to asking an AI agent instead of searching for themselves. If your catalog isn't reachable, it isn't part of the conversation. That's not because it's worse. It's because it's unreachable.

Enter the Model Context Protocol

Nobody has to solve this alone, from scratch, with a custom integration for every AI tool. There's now an open standard for it: the Model Context Protocol, or MCP. In plain terms, MCP lets an AI tool call into an external system through a small set of well-defined, typed actions, called tools, instead of scraping a website or guessing from a product description. A client like Claude Desktop or Cursor can connect to any MCP server and immediately know what actions are available and what each one expects, the same way a plugin declares its inputs and outputs.

That's the door. On its own, it doesn't put a catalog behind it.

What we're building: sound-catalog-mcp

That's what sound-catalog-mcp is for: a self-hostable MCP server that exposes a sound catalog to any MCP-speaking agent. It's read-only by design; no agent can upload, delete, or edit anything through it. It runs entirely on infrastructure you control, with no third-party AI API calls required at query time. And it's generic. Any catalog directory that follows its schema runs on the same codebase. The Insect Audio catalog is the first real instance of it, not a special case wired into the code.

Point an agent at it, and a question like "find me something metallic and low-pitched I can use commercially" stops being a dead end. The agent can search the catalog by meaning instead of just keyword, pull up exactly where a sound came from and how it was made, check its license instead of guessing, and hand back a link to actually listen, all without your catalog data ever leaving infrastructure you control.

Where this series goes next

This is the first of five articles walking through how it's built, each one going a level deeper than the last:

  1. This article: why an invisible catalog is a real problem, not a hypothetical one.
  2. Provenance first: why a sound's backstory and licensing terms belong in the core data model, not bolted on as an afterthought.
  3. Search, without a cloud API call: how the server finds the right sound by meaning, entirely on local infrastructure.
  4. Five tools, zero write access: the protocol layer, and the security model that makes it safe to hand an agent direct access.
  5. Run it yourself: a hands-on walkthrough, from an empty folder to a working, self-hosted server.

If you'd rather not wait for the rest of the series, the project is open source right now, at github.com/insectaudio/sound-catalog-mcp.

Tags