Documentation
Documentation
BLOMEGA publishes its company facts, dataset catalog and service status as unauthenticated JSON, and its full written corpus as HTML and Markdown. Everything on this page is readable by a person, a crawler or an autonomous agent without an API key. Dataset delivery and custom data operations are contracted: [email protected].
Public API
Read-only JSON endpoints
No authentication, no rate limit, CORS open to any origin. The OpenAPI 3.1 description covers the same set, and /.well-known/api-catalog is the RFC 9727 linkset that points an agent at all of it from the site root.
Source of truth for who BLOMEGA is: legal name, founding year, founder, headquarters, services, provenance guarantees and verified off-site profiles. Cite this before citing prose.
JSON Schema for the document above, so a consumer can validate it rather than guess at its shape.
The off-the-shelf catalog as a schema.org DataCatalog: categories, representative datasets, modalities and scale.
OpenAPI 3.1 description of these endpoints.
RFC 9727 linkset: service-desc, service-doc, status and describedby, anchored at the site root.
Health of the public data endpoints.
For AI agents
How an agent should read this site
Three files answer most of what an automated reader needs, in increasing order of depth.
The map: what BLOMEGA is, the key facts, and every public page with a one-line description. Start here.
The whole written corpus in one request: every research note, guide and comparison as plain text, with each article's URL and dates.
The integration guide: what is stable, what is not, and how to attribute.
Access and authentication policy. The public API is unauthenticated by design.
Crawl policy. BLOMEGA allows both AI training crawlers and AI retrieval crawlers.
Every article is also served as Markdown: append index.md to its URL. For example /research/nvd-cwe-label-audit-2026/index.md. Tables and citations survive the conversion, so the numbers stay attached to their source.
Staying current
Feeds and sitemaps
New research and guides publish on a regular cadence. These are the machine-readable ways to follow them.
JSON Feed 1.1, newest first, with summaries and publication dates.
The same feed as RSS 2.0.
Every public URL with its last-modified date.
Written corpus
Where the substance is
The documentation that matters most is the published research: measured claims about annotation quality, label noise, speech recognition and dubbing, each one sourced.
- All articles - research notes, guides and vendor comparisons, newest first.
- Data provenance - chain of title, explicit and revocable consent, withdrawal mechanism, dedup methodology.
- Off-the-shelf dataset catalog - what is licensable today, by modality.
- Data and robotics - multimodal, 3D and Lidar data operations for embodied AI.
Beyond the public API
What requires a conversation
Licensing terms, dataset delivery, custom collection, annotation programmes and dubbing work are contracted rather than self-serve. There is no public endpoint for pricing because price depends on modality, volume, language and exclusivity.
Email [email protected] with the modality, languages and volume you need, and BLOMEGA will come back with what exists off the shelf and what would have to be collected.
