LICENSED REAL-WORLD DATA FOR AI

Real-world data.
Precisely
prepared.

Audio, video, and creator content for AI teams.
Clear specifications. Documented rights.
A delivery built around your requirements.

Human at the source. Precise at delivery.
Many layers. One considered dataset.Select a layer to take a closer look.
BUILT AROUND YOUR BRIEF
Audio & speechVideo & UGCMultimodal collectionsCustom capture
BUILT AROUND YOUR MODEL

The details define
the dataset.

Start with the task. Work back to the source.

COLLECTION DESIGN / 01
An illustration of collection requirements
Speech recognition & diarization

Speech with the texture
of real conversation.

Specify the speakers, languages, overlap, and recording environments that matter to your task.

Source
Natural conversation
Conditions
Speakers · accents · environment
Annotations
Transcripts · turns · timestamps
Start a speech brief

Opens an email with a starting outline. Availability and final scope are confirmed together.

AN INTERACTIVE LOOK INSIDE

One example.
Every layer connected.

A file is only the beginning. Explore how media, annotations, and documentation can come together in a usable dataset.

Anatomy of an aligned record.
Illustrative example
EXAMPLE / 001
Aligned media: Speaker A, 00:04.200 to 00:08.600
00:04.200ILLUSTRATION
TIME-ALIGNED ANNOTATIONS01—03
00:04.20000:18.400
Select a segment or inspect the fields.Synthetic example · Deliverables vary by collection
FROM THE FILE TO THE FIELD

Good data should be
easy to examine.

A considered handoff includes the structure around the media. Explore an example of the fields your engineering team can inspect.

example_conversation_001

Specifications that travel with the files.

Resolution, frame rate, sample rate, and channel count belong in the delivery conversation.

Explore the diligence package
manifest.jsonSYNTHETIC EXAMPLE
example_conversation_001/ media
{  "video": {    "width": 1920,    "height": 1080,    "frame_rate": 25  },  "audio": {    "sample_rate_hz": 48000,    "channels": 2  }}
Download the complete example

Synthetic data, shown to explain the structure. Actual fields and deliverables are scoped per collection.

Review rights & provenance
THE PATH TO DELIVERY

Care in every layer.

Technical fit and clear permissions belong in the same conversation. Here's how a brief becomes a defined delivery.

FROM BRIEF TO DELIVERY01 / 03
SOURCE
Input
Your dataset brief
Review
Source & task fit
Output
A scoped collection

An illustration of the process. Scope is agreed per project.

FOR THE PEOPLE WHO MADE IT

Your work has value.
Keep ownership of it.

Give your audio, video, or creator library a new opportunity through AI licensing. We help connect your work with relevant buyers and document the scope.

Explore licensing your library
A FEW GOOD QUESTIONS

Clear from
the beginning.

For AI teams and the people
who own the media.

More in our resources
How do AI teams license audio and video data?

You send a spec: modality, language, hours, recording conditions. We match it against libraries whose owners have agreed to license, clear the rights, and deliver the data under a licence that grants AI-training use explicitly. One agreement covers the whole delivery.

What kinds of content can be licensed for AI training?

Original audio, video, creator UGC, motion capture, sensor data, task demonstrations, and other media may be licensable when the supplier controls the relevant rights and the proposed AI use, identifiable people, embedded works, privacy, and contractual restrictions are addressed. A category page is not an availability claim; a particular collection still has to pass review.

How do data owners get paid?

Per licence. You keep ownership while fiund handles AI licensing under one contributor agreement, and you are paid each time an AI team licenses your material. What a deal is worth depends on modality, exclusivity and volume — the payouts guide explains how deals are structured.

How does fiund verify provenance and consent?

We connect each approved asset to its source, owner or authorized licensor, applicable agreement, consent status, known restrictions, and file-level manifest. Unresolved material is held back instead of being described as cleared.

How much does it cost to license training data?

There is no universal rate card. Pricing is scoped to the content type, task fit, volume, source quality, metadata, consent and diligence work, permitted uses, delivery requirements, and any requested exclusivity. A quote follows the actual brief and verified inventory.

What documentation arrives with licensed data?

An approved delivery can include the buyer licence, a versioned file manifest, data dictionary, technical specifications, provenance and consent references, known restrictions, and integrity evidence for the exact transferred set.

More about licensing and permitted use
What is AI training data licensing?

A signed agreement that lets an AI team train models on media someone else owns — audio, video, images or text — in exchange for payment. The licence spells out what can be trained, for how long, and what rights travel with the data, so nobody is relying on fair-use guesses.

Can an AI team buy training data with commercial rights?

Yes, through a written buyer licence. The agreement should identify the dataset, approved AI purposes, production or evaluation scope, recipients, term, territory, retention, and restrictions. A download, public post, or evaluation sample is not automatically a production-training grant.

Is scraped data safe to train on?

Public accessibility alone does not establish permission for a proposed AI use. Buyers should examine the source, applicable permissions, privacy considerations, restrictions, and intended use. Our rights guides explain the questions to raise during diligence.

What does rights-cleared mean?

Every permission needed for AI training is secured in writing before data changes hands: ownership confirmed, training rights granted explicitly in the licence, and voice and likeness consent covered wherever people are identifiable.

Does fiund authorize voice or likeness cloning?

No. fiund’s standard contributor terms do not authorize a buyer to create a synthetic replica, clone, or voice or likeness model intended to imitate an identifiable person, or use that identity to imply endorsement.

How quickly can fiund respond to a dataset brief?

fiund targets an initial scope and availability response within two business days. That response distinguishes current inventory from data that would need to be sourced or commissioned; it is not a promise that every collection can be delivered in that window.

LET'S START WITH YOUR BRIEF

What does your
next model need?

Tell us the modality, the task, and the conditions that matter. We'll help define the next step.

Discuss your dataset jaeden@fiund.com