Real-world data.
Precisely
prepared.
Audio, video, and creator content for AI teams.
Clear specifications. Documented rights.
A delivery built around your requirements.
For models that understand
more of the real world.
Speech, video, and the connections between them. Explore the kinds of data your next model could learn from.
Human conversation.
Every dimension of it.
Accents, speaker turns, interruptions, and the conditions beyond the studio.
Context in motion.
Footage, sequences, and source context for models that work with the visual world.
Connected signals.
A richer picture.
Define how audio, video, text, and metadata need to fit together.
The world,
from another viewpoint.
Explore task demonstrations, first-person perspectives, and continuous action.
A particular task.
A considered collection.
Scope the recording conditions, participants, and outputs your brief calls for.
Illustrative graphics. Collection availability, permissions, and custom-capture feasibility are confirmed against your brief.
The details define
the dataset.
Start with the task. Work back to the source.
Speech with the texture
of real conversation.
Specify the speakers, languages, overlap, and recording environments that matter to your task.
- Source
- Natural conversation
- Conditions
- Speakers · accents · environment
- Annotations
- Transcripts · turns · timestamps
Opens an email with a starting outline. Availability and final scope are confirmed together.
One example.
Every layer connected.
A file is only the beginning. Explore how media, annotations, and documentation can come together in a usable dataset.
Explore the example. Then let's talk about your actual requirements.
Explore dataset guidesGood data should be
easy to examine.
A considered handoff includes the structure around the media. Explore an example of the fields your engineering team can inspect.
Specifications that travel with the files.
Resolution, frame rate, sample rate, and channel count belong in the delivery conversation.
{ "video": { "width": 1920, "height": 1080, "frame_rate": 25 }, "audio": { "sample_rate_hz": 48000, "channels": 2 }}Download the complete example Care in every layer.
Technical fit and clear permissions belong in the same conversation. Here's how a brief becomes a defined delivery.
- Input
- Your dataset brief
- Review
- Source & task fit
- Output
- A scoped collection
An illustration of the process. Scope is agreed per project.
Permissions at the source.
Clarity at the handoff.
Understand the source, the scope, and the handoff.
Get specific before you commit.
Know the source.
Review how ownership, permissions, and consent connect to the media.
Rights & provenanceUnderstand the scope.
Inspect permitted use, known restrictions, and the documentation available for diligence.
The diligence packageDefine the delivery.
Agree specifications, formats, annotations, and the records that accompany the files.
Working with FIUNDYour work has value.
Keep ownership of it.
Give your audio, video, or creator library a new opportunity through AI licensing. We help connect your work with relevant buyers and document the scope.
Explore licensing your libraryThink deeper. Build better.
What rights-cleared really means.
The permissions and documentation to examine before a collection changes hands.
Read the guide THE BUYER'S GUIDEStart with the right questions.
Make the proposed use, scope, and delivery requirements explicit from the beginning.
Read the guide CREATOR CONTENTA new context for original work.
Explore how creator content can be licensed for AI, and what needs to be agreed.
Read the guideClear from
the beginning.
For AI teams and the people
who own the media.
How do AI teams license audio and video data?
You send a spec: modality, language, hours, recording conditions. We match it against libraries whose owners have agreed to license, clear the rights, and deliver the data under a licence that grants AI-training use explicitly. One agreement covers the whole delivery.
What kinds of content can be licensed for AI training?
Original audio, video, creator UGC, motion capture, sensor data, task demonstrations, and other media may be licensable when the supplier controls the relevant rights and the proposed AI use, identifiable people, embedded works, privacy, and contractual restrictions are addressed. A category page is not an availability claim; a particular collection still has to pass review.
How do data owners get paid?
Per licence. You keep ownership while fiund handles AI licensing under one contributor agreement, and you are paid each time an AI team licenses your material. What a deal is worth depends on modality, exclusivity and volume — the payouts guide explains how deals are structured.
How does fiund verify provenance and consent?
We connect each approved asset to its source, owner or authorized licensor, applicable agreement, consent status, known restrictions, and file-level manifest. Unresolved material is held back instead of being described as cleared.
How much does it cost to license training data?
There is no universal rate card. Pricing is scoped to the content type, task fit, volume, source quality, metadata, consent and diligence work, permitted uses, delivery requirements, and any requested exclusivity. A quote follows the actual brief and verified inventory.
What documentation arrives with licensed data?
An approved delivery can include the buyer licence, a versioned file manifest, data dictionary, technical specifications, provenance and consent references, known restrictions, and integrity evidence for the exact transferred set.
More about licensing and permitted use
What is AI training data licensing?
A signed agreement that lets an AI team train models on media someone else owns — audio, video, images or text — in exchange for payment. The licence spells out what can be trained, for how long, and what rights travel with the data, so nobody is relying on fair-use guesses.
Can an AI team buy training data with commercial rights?
Yes, through a written buyer licence. The agreement should identify the dataset, approved AI purposes, production or evaluation scope, recipients, term, territory, retention, and restrictions. A download, public post, or evaluation sample is not automatically a production-training grant.
Is scraped data safe to train on?
Public accessibility alone does not establish permission for a proposed AI use. Buyers should examine the source, applicable permissions, privacy considerations, restrictions, and intended use. Our rights guides explain the questions to raise during diligence.
What does rights-cleared mean?
Every permission needed for AI training is secured in writing before data changes hands: ownership confirmed, training rights granted explicitly in the licence, and voice and likeness consent covered wherever people are identifiable.
Does fiund authorize voice or likeness cloning?
No. fiund’s standard contributor terms do not authorize a buyer to create a synthetic replica, clone, or voice or likeness model intended to imitate an identifiable person, or use that identity to imply endorsement.
How quickly can fiund respond to a dataset brief?
fiund targets an initial scope and availability response within two business days. That response distinguishes current inventory from data that would need to be sourced or commissioned; it is not a promise that every collection can be delivered in that window.
What does your
next model need?
Tell us the modality, the task, and the conditions that matter. We'll help define the next step.
Discuss your dataset jaeden@fiund.com