---
title: "AI Model Pre-Training Optimization: Raw Data | KUZOG Ventures"
description: "KUZOG Ventures research on AI model pre-training optimization: how companies acquire raw training data, at what cost, and how to get more value per token."
canonical: "https://www.kuzog.com/ventures/research/ai-pre-training-data-acquisition/"
last_updated: 2026-10-02
language: en
---

# AI Model Pre-Training Optimization: Raw Data | KUZOG Ventures

> KUZOG Ventures research on AI model pre-training optimization: how companies acquire raw training data, at what cost, and how to get more value per token.

Canonical: https://www.kuzog.com/ventures/research/ai-pre-training-data-acquisition/
Language: en

A model learns most of what it knows in pre-training, and pre-training is only as good as the raw data that goes in. We work with two companies on one question: how to acquire better raw data for pre-training, at a lower cost.

Status: Ongoing research with two industry partners. No results published yet.

## Why raw data is the bottleneck

Compute gets cheaper every year. Good raw data does not. The open web is crawled many times over, much of it is duplicated or low quality, and the rights to use it are less clear than they were.

So the cost of a pre-training run moves from the GPUs to the data pipeline: finding sources, getting the right to use them, cleaning them, and proving what is inside.

## What we work on

With our two partners, we study each step from a raw source to a training set. The aim is more useful tokens for each euro spent, and a clear record of where every token came from.

- Sourcing: which kinds of raw data add new knowledge, and which only repeat what the model has already seen.
- Rights and provenance: how to acquire data with clear terms, and keep a record a buyer can audit.
- Filtering and deduplication: how much to remove, and when cleaning starts to remove signal.
- Value per token: how to estimate what a dataset adds before paying for a full training run.

## Who this is for

Teams that train or fine-tune their own models, and companies that hold data and want to supply it to them. If you work on either side and have a question for this project, send us a pitch.

## Pitch us a research idea

Do you have a question you want to research with us? Send a short pitch: the problem, why it matters, and what you bring.

Contact: https://www.kuzog.com/contact/

## Sitemap

- [Sitemap (Markdown)](https://www.kuzog.com/sitemap.md): every page in every language, with its Markdown twin.
- [sitemap.xml](https://www.kuzog.com/sitemap.xml)
- [llms.txt](https://www.kuzog.com/llms.txt)

---

_Source: https://www.kuzog.com/ventures/research/ai-pre-training-data-acquisition/ · Markdown variant served via Accept: text/markdown · Full brief: https://www.kuzog.com/llms.txt_
