Find the data. Fund the people who made it.
A licensing marketplace for real-world audio, video and creator content. AI teams get training data that isn't already in the crawl. The people who own it get paid, and keep their rights.
Real-world media, sourced from the people who own it.
Audio, video and creator UGC that never made it into a public crawl. And if the dataset you need doesn't exist yet, we capture it to spec.
Conversational audio
Real people talking to each other — interruptions, crosstalk, accents. The other kind of speech data, the kind models are short on.
Film & real-world video
Natural rooms, natural light, real motion.
Creator libraries
Short-form, phone-shot, self-filmed — at influencer scale.
Custom capture
Doesn't exist yet? We film it to spec, rights signed on set.
Egocentric · motion capture · sensor & IMU · agentic trajectories
Every licence flows through the middle.
Owners keep their rights. AI teams get clean, defensible data. fiund sits in the middle and holds the paperwork.
Licensed at the source. Every single asset.
Training-data deals fall apart over provenance, not quality. Every row here is backed by a signed agreement you can pull in diligence.
Questions, answered plainly.
The things AI teams and data owners ask us most — without the legalese.
What is AI training data licensing?
A signed agreement that lets an AI team train models on media someone else owns — audio, video, images or text — in exchange for payment. The licence spells out what can be trained, for how long, and what rights travel with the data, so nobody is relying on fair-use guesses.
How do AI teams license audio and video data?
You send a spec: modality, language, hours, recording conditions. We match it against libraries whose owners have agreed to license, clear the rights, and deliver the data under a licence that grants AI-training use explicitly. One agreement covers the whole delivery.
What kinds of content can be licensed for AI training?
Original audio, video, creator UGC, motion capture, sensor data, task demonstrations, and other media may be licensable when the supplier controls the relevant rights and the proposed AI use, identifiable people, embedded works, privacy, and contractual restrictions are addressed. A category page is not an availability claim; a particular collection still has to pass review.
Can an AI team buy training data with commercial rights?
Yes, through a written buyer licence. The agreement should identify the dataset, approved AI purposes, production or evaluation scope, recipients, term, territory, retention, and restrictions. A download, public post, or evaluation sample is not automatically a production-training grant.
Is scraped data safe to train on?
It is a risk you inherit forever. Scraped material arrives with no rights, no consent and no chain of title — which is where most training-data disputes start. Our rights guides and the licensed-vs-scraped comparison below walk through the difference in detail.
How do data owners get paid?
Per licence. You keep ownership while fiund handles AI licensing under one contributor agreement, and you are paid each time an AI team licenses your material. What a deal is worth depends on modality, exclusivity and volume — the payouts guide explains how deals are structured.
What does rights-cleared mean?
Every permission needed for AI training is secured in writing before data changes hands: ownership confirmed, training rights granted explicitly in the licence, and voice and likeness consent covered wherever people are identifiable.
How does fiund verify provenance and consent?
We connect each approved asset to its source, owner or authorized licensor, applicable agreement, consent status, known restrictions, and file-level manifest. Unresolved material is held back instead of being described as cleared.
Does fiund authorize voice or likeness cloning?
No. fiund’s standard contributor terms do not authorize a buyer to create a synthetic replica, clone, or voice or likeness model intended to imitate an identifiable person, or use that identity to imply endorsement.
How much does it cost to license training data?
There is no universal rate card. Pricing is scoped to the content type, task fit, volume, source quality, metadata, consent and diligence work, permitted uses, delivery requirements, and any requested exclusivity. A quote follows the actual brief and verified inventory.
How quickly can fiund respond to a dataset brief?
fiund targets an initial scope and availability response within two business days. That response distinguishes current inventory from data that would need to be sourced or commissioned; it is not a promise that every collection can be delivered in that window.
What documentation arrives with licensed data?
An approved delivery can include the buyer licence, a versioned file manifest, data dictionary, technical specifications, provenance and consent references, known restrictions, and integrity evidence for the exact transferred set.
Let's talk about what you actually need.
Whether you're building a model or sitting on an archive, the first conversation is short and specific.
Get in touch