A licensing marketplace for real-world audio and video. AI teams get training data that isn't already in the crawl. The people who own it get paid, and keep their rights.
We focus on material that never made it into a public crawl — recorded by real people, in real conditions, and licensed directly from them.
Most speech corpora are scripted and read aloud. We source the other kind — people actually talking to each other, thinking out loud, talking over one another.
Every asset carries a signed licence and, where people are identifiable, voice and likeness consent. The paperwork travels with the file.
Creators, studios and archives on the supply side — approving every use, keeping ownership, and getting paid per licence.
Volume gets you a demo. Chain of title gets you a model you can ship.
Non-public material sourced directly from owners, so you aren't paying for something your model has already seen.
Every asset carries a signed licence from its owner granting AI training rights explicitly. We can show you the paperwork.
If the dataset doesn't exist yet, tell us the spec and we go source it. That's the part the name is about.
If you've been recording for years, you're sitting on an asset. Licensing it shouldn't mean losing it.
You license, you don't sell. The material stays yours and you set the terms it goes out under.
Nothing is licensed to a buyer or a use case you haven't signed off on.
Transparent rates, transparent terms, and no exclusivity lock-in unless you want it and it's priced for it.
A buyer sends a brief — modality, languages, conditions, volume, licence terms. Or an owner sends us what they have and we tell them honestly whether there's a market for it.
We go to the people who hold the material and paper the rights before anything moves. Consent, likeness and training rights are handled up front, not retroactively.
The dataset arrives with its documentation attached. The owner gets paid. Everyone can see exactly what was licensed and to whom.
Training data deals almost never fall apart over quality. They fall apart because nobody can prove where the material came from or what rights came with it. We built around that problem first.
A written agreement with the owner of every asset we license — no exceptions, no assumptions.
Granted in the licence itself rather than inferred from a general content or distribution agreement.
Covered separately wherever people are identifiable, because copyright alone doesn't reach it.
No crawling, and nothing sourced from behind a login. Provenance records are available during diligence.
Whether you're building a model or sitting on an archive, the first conversation is short and specific.