$5m · Other round · Software · London, UK
Mozilla Data Collective, a British social enterprise providing vetted and licensed datasets for AI development, has raised $5m in an investment from Mozilla. The company says the funding will expand its catalogue with multimodal cultural video and larger text corpora, broaden licensing and subscription options, and strengthen platform security and data-improvement tools.
Licensed data turns provenance into product infrastructure
AI teams need more than volume when training across languages and media: they need permissioned material, documented provenance and terms that survive procurement scrutiny. Mozilla Data Collective's catalogue model packages those requirements into a searchable supply layer. Its reported breadth across more than 1,700 datasets and 450-plus languages gives buyers a route beyond repeatedly negotiating with individual data holders.
The commercial test is whether stronger licensing options and subscriptions can make that supply dependable for both sides. Better curation and security can reduce integration and compliance work for model builders, while recurring purchase structures can give contributors clearer incentives to keep specialised datasets available. The round therefore finances not just catalogue growth, but the trust and delivery mechanics required for licensed data to become reusable infrastructure.


