BLO Mega

Data Provenance

Chain of title.
Consent on record.
Audit on demand.

Every dataset we deliver is manufactured to spec — not scraped, not brokered. This page answers the due-diligence questions serious buyers now send before signing.

Chain of Title

Every asset traceable to the originating contributor, contract, and delivery batch.

Speaker & Subject Consent

Signed, versioned consent captured before collection — stored with the asset.

Withdrawal Mechanism

Contributors can revoke; assets are pulled from active and derivative datasets on request.

Dedup Methodology

Perceptual hashing, embedding similarity, and manual review to eliminate leakage.

01 — Chain of Title

From contributor to delivery, one unbroken record.

Each contribution is bound at capture time to a contributor ID, a signed contributor agreement (versioned), a project brief, and a delivery batch. That binding travels with the asset through annotation, QA, and export — the manifest we deliver alongside every dataset reproduces it row by row.

No third-party brokered corpora enter our production pipelines. Where a client supplies upstream material, we require documented rights and reflect that upstream provenance in the manifest.

02 — Speaker & Subject Consent

Explicit, informed, revocable.

Voice contributors, on-camera subjects, and document donors sign a consent that names the intended use (AI training, evaluation, derivative model deployment) in plain language, and identifies the contracting entity. Consent is captured before collection begins and re-affirmed for any material change of scope.

Minors, protected classes, and sensitive-content collections follow additional review — including guardian consent where applicable — and are flagged in the manifest so downstream teams can honor use restrictions.

03 — Withdrawal Mechanism

Contributors can leave. Their data leaves with them.

A contributor can request withdrawal through a documented channel. On confirmation, their assets are removed from active datasets, quarantined from future exports, and reflected in a withdrawal ledger delivered to affected clients within the contracted SLA. Derivative datasets built from withdrawn material are re-issued without it.

Withdrawal does not retroactively invalidate models already trained on prior versions — that boundary is stated explicitly to contributors at consent time.

04 — Dedup Methodology

No leakage between train, eval, and holdout.

Deduplication runs at three layers: exact-hash for identical assets, perceptual / embedding similarity for near-duplicates (audio fingerprinting, image and text embeddings), and human adjudication on flagged clusters. Split assignments are frozen after dedup so that eval and holdout sets remain uncontaminated across re-issues.

Cross-batch and cross-project dedup are available for clients running long-lived training pipelines against multiple deliveries.

05 — Handling, Storage, Access

Least privilege, logged access, encrypted at rest and in transit.

Contributor PII is segregated from training assets and available only to a named operations team. Client-facing exports contain the training data plus an anonymized manifest; identity mapping is retained internally for audit and withdrawal handling.

Access to production data requires named accounts, MFA, and role-scoped permissions. Access logs are retained and available under NDA for buyer due diligence.

Buyer Due Diligence

Answers to the questionnaire you're about to send.

Is any of this data scraped from the public web?

No. Our delivery pipeline produces manufactured, consented data. Web-sourced material only enters as reference under a documented rights basis provided by the client.

Do you own the delivered assets or license them from brokers?

Neither. We produce the data under contributor agreements that assign the rights the client contracts for. No brokered corpora sit in the delivery path.

How do you handle contributor withdrawal after delivery?

Withdrawal requests are honored on a contracted SLA. Assets are removed from active sets and re-issued datasets; a withdrawal ledger is delivered to affected clients.

What prevents train / eval / holdout leakage across releases?

Three-layer dedup (exact hash, embedding similarity, human adjudication) with frozen split assignments and optional cross-batch dedup for multi-release clients.

Can we audit consent for any individual asset?

Yes. Each manifest row references a contributor ID and consent version; under NDA we can produce the underlying signed record for named assets.

How is contributor PII stored and separated from training data?

PII is stored in a separate identity system with role-scoped access. Training exports contain anonymized manifests; identity mapping is retained internally for audit and withdrawal.

What is your position on synthetic augmentation?

Synthetic augmentation is available on request and clearly flagged in the manifest so it can be excluded from eval or holdout sets.

Send the questionnaire. We'll send the manifest.

Provenance packet available under NDA — contributor agreements, consent samples, dedup report, and a walk-through of a live delivery manifest.

Reviewed quarterly · Versioned · Client-auditable