Photo by Markus Winkler / Unsplash

Provenance First: Why We Put a Sound's Backstory in the Data Model (2)

provenance Aug 8, 2026

The first article in this series was about a door: giving an AI agent a way into a sound catalog at all, through the Model Context Protocol. This one is about what's actually behind that door. Because opening a door onto a mess doesn't help anyone. What matters is the shape of the data the agent finds once it's in, and that shape is where most of the real design decisions live.

The failure mode we're avoiding

Ask a language model to summarize a license, and it will happily oblige, whether or not it actually understands the license. That's the risk. "You can use this for non-commercial projects with attribution" can drift, three conversational turns later, into "you can use this." Nobody lied. The model just did what language models do: filled in a plausible-sounding gap. In a casual conversation that's a minor annoyance. In a legal claim about what someone can and can't do with a piece of audio, it's a real problem, and it's one a catalog owner has no visibility into once it happens inside someone else's chat window.

So the question we started from wasn't "how do we describe a sound to an AI agent." It was "how do we make it structurally difficult for an AI agent to say something false about a sound," and that turns out to be mostly a data modeling problem, not a prompting problem.

Five facets, not one metadata blob

Most sample libraries describe what a sound is: a title, some tags, maybe a category. That's usually where it stops. Our schema treats the story of a sound as just as important as the sound itself, and it splits that story into five named sections that mirror how a sound actually gets made, rather than dumping everything into one free-text field:

  • source_object, the thing that produced the sound and how it was excited, whether that's a physical object being struck or bowed, a field recording of something happening in the world, or (as in our own demo catalog) a synthesis algorithm.
  • capture, the recording chain, or its digital equivalent, including space, date, and sample rate.
  • processing, what was done to the raw sound afterward, from normalization to heavier editing.
  • license, exactly what a user may and may not do with the sound, expressed as structured fields rather than a paragraph of legal prose.
  • integrity, whether any part of the sound involved AI generation, and whether the audio file still matches what the catalog claims it is.

None of this is exotic. It's closer to how a sound designer already thinks about a sample than to how a database schema usually gets built. The point of writing it down this way is that a search layer can now filter on license.commercial_use or integrity.ai_generated as real fields, and an agent asked to describe a sound has actual structured facts to draw from instead of a description string it has to interpret and hope it got right.

What this looks like in practice

Here's a real entry from the catalog bundled with the project, trimmed to the fields that matter for this article:

{
  "id": "design-granular-texture-01",
  "title": "Evolving granular noise texture",
  "source_object": {
    "type": "granular_noise_synthesis",
    "material": "~220 windowed, individually filtered white-noise grains"
  },
  "capture": {
    "space": "digital synthesis, no physical recording space",
    "sample_rate": 48000
  },
  "processing": ["normalize"],
  "license": {
    "type": "cc-by-4.0",
    "commercial_use": true,
    "attribution_required": true,
    "redistribution": true
  },
  "integrity": {
    "ai_generated": false,
    "synthetic_content": "Fully synthesized via granular synthesis over filtered white-noise grains; no recorded source audio."
  }
}

A few things worth noticing. source_object and capture together tell you this sound was built entirely in software, not recorded in a room; there's no microphone or space to describe, so the schema says so plainly instead of leaving those fields awkwardly blank. license answers "can I use this commercially, and do I need to credit someone" without a sentence of legal writing anywhere near it. And integrity.ai_generated is false even though the sound is fully synthesized, which is a distinction worth sitting with for a second: synthetic doesn't mean AI-generated, and a catalog that can't tell an agent the difference is quietly setting it up to get that wrong.

The rule that makes the license field trustworthy

Structured fields alone don't fully solve the paraphrasing problem. Something still has to turn commercial_use: true and attribution_required: true into words a person or an agent can read, and that translation step is exactly where over-claiming creeps back in if you're not careful about it.

The rule we hold that translation to is this one, and it's the single most important sentence in this article:

Never paraphrase legal text beyond what the catalog states.

In practice that means the code that builds a license summary is not allowed to branch on the license type at all. It only ever assembles fixed sentence fragments from the three or four boolean fields the catalog actually states. It can't say "public domain" for a CC0 sound, because "public domain" isn't quite the same thing in every jurisdiction and the catalog never claimed that word. It can only say what the booleans say. That's not a policy someone has to remember to follow. It's mechanically true of the code, because the function has no other information available to reach for.

The full legal text, when a catalog has any, lives as a plain text file next to the catalog itself and gets read at request time rather than baked into the tool's own logic. That keeps the actual legal wording under the catalog owner's control, not ours.

One more piece: proving the file didn't quietly change

The fifth facet, integrity, does one more thing worth a mention: it stores a hash of the audio file itself. That sounds like a minor detail until you picture the alternative, an agent describing a sound and handing back a license summary for a file that was silently swapped out or corrupted at some point after the catalog was written. The hash means the system can tell, and refuse to serve stale or mismatched data instead of confidently describing the wrong thing.

All of this, incidentally, is a published, versioned JSON Schema, not just an internal convention. A catalog owner (or eventually, a third party) can validate a catalog.json against it independently of running any of our code. More on that when this series gets to actually running the project yourself.

None of it matters yet, though, if an agent can't find the right sound among a few thousand entries in the first place. That's the next article: how search works, and how it runs without sending a single query to a third-party AI API.

Tags