# Partner with us | Ooak Data

Y Backed by Y Combinator 

# Turn your data into revenue

Ooak Data licenses the operational and business data that defines how companies actually work, anonymizes it, and turns it into training material for frontier AI models. You get paid for something you already have.

- Earn up to $300K
 - 2 to 4 weeks
 - No engineering work required

 [Book a call](https://calendar.google.com/appointments/schedules/AcZssZ142jxQGEjB1PR9skAe17Nf7Pjf8G8EszSPcRhZIzmZXGCkn05REmmAmqPvMgWTYUMwbQNwinIF) [What is my data worth? ](#estimator) 

Estimator

## What is my data worth?

Connected data is worth more than big data. Every tool you include adds another thread to the story, and the value is in the threads that join up.

 What is my data worth? 

Answer a few questions for a preliminary estimate.

From your systems to the model

## The data that teaches models how work actually gets done

Every tool your company ran on, anonymized by Ooak into one consistent dataset.

- Emails
- Slack & Teams
- Documents
- Tickets
- Code & PRs
- CRM
- Support
- Spreadsheets

- Intelligence
- Efficiency
- Velocity
- Reliability
- Autonomy
- Judgment

  Anonymized 

Your systems

What models gain

Why now

## The AI labs need new sources of data

1. 01

### The easy supply is exhausted

Web text, books, code and forums have all been scraped and trained on.

 2. 02

### Models now need complexity

Longer, multi-step, multi-tool tasks take richer datasets, not more short-form text.

 

A single workflow, across multiple systems

1. Customer email
 2. 
 3. Support ticket
 4. 
 5. Slack thread
 6. 
 7. Jira issue
 8. 
 9. Pull request
 10. 
 11. CRM update

[Read our research →](/research)

Your part

## Licensing your company data

We run the whole process, from the first call to the signed agreement. The technical work is ours.

Book a call and tell us which tools you are ready to license.

 [Book a call](https://calendar.google.com/appointments/schedules/AcZssZ142jxQGEjB1PR9skAe17Nf7Pjf8G8EszSPcRhZIzmZXGCkn05REmmAmqPvMgWTYUMwbQNwinIF) 

1.  1 

### Chat with our team

We map your workflows, tools and data footprint, and give you an estimate of what your data is worth. Fifteen minutes.

 2.  2 

### Decide what is in scope

You pick the systems and date ranges. Anything you exclude is annexed to the contract before a single record moves.

 3.  3 

### We evaluate your data

Hand in hand, we count what is actually there: volume, coverage and how deeply your tools connect. Under an hour of your time, and not a line of engineering work.

 4.  4 

### We anonymize your data

Our in-house engine strips all PII. Names, companies and addresses are replaced consistently across every source, and faces, logos and signatures are blacked out.

 5. 5 

### Get paid

We pay fast, and your raw archive is deleted within 30 days - a contract obligation, not a policy.

 

Our core technology

## Anonymization is not a step in our process. It is what we build.

We built the engine, we run it, and no third party ever touches your raw archive. One entity dictionary is applied across every connected source, so a name replaced in an email is the same replacement in the ticket, the ledger and the commit. Faces, logos and signatures are detected and blacked out, page by page.

Your source

 #launch-hydra-serum Camille Roux 09:12 

@Théo we are 4,000 units short for the Vantelle FR drop on the 18th. Milan plant confirmed 11,000 of 15,000.

Entity dictionary

- Camille Roux→Sofia Klein
 - Théo→Marc
 - Vantelle FR→Retailer 3
 - Milan→City 2

One dictionary, every source. Same input, same output.

Anonymized

 #launch-product-a Sofia Klein 09:12 

@Marc we are 4,000 units short for the Retailer 3 drop on the 18th. City 2 plant confirmed 11,000 of 15,000.

The workflow survives. The identities do not. That consistency across sources is what makes an anonymized dataset usable, and it is the hard part - we master it.

What a dataset looks like

## Three datasets, three shapes

Three of the datasets we have licensed. Different sectors, different sizes, different stacks.

- 

### Edtech

 People400+History6 yearsTools15 

Google Workspace, MySQL, BigQuery, MongoDB, GitLab, ClickUp, AWS, Azure, GCP

 - 

### Construction tech

 People200+History4 yearsTools7 

Google Workspace, Slack, Jira, Asana, Airtable, GitHub, HubSpot

 - 

### Fleet logistics

 People15+History4 yearsTools7 

Google Workspace, Slack, WhatsApp, Jira, Notion, Supabase, GitHub

 

FAQ

## Questions founders ask

- 

### Am I allowed to sell this data? 

In most cases, yes. Your company owns its operational records, and the contract licenses them for training only, with your exclusions annexed. We check jurisdiction and requirements during scoping. Counsel reviews everything before signature.

- 

### Can a buyer identify my company? 

The buyer receives a cleaned dataset under a code name. Company names, people, addresses and logos are replaced or blacked out consistently across every source. We have recorded zero re-identification incidents to date. You review a sample before delivery.

- 

### Who actually sees my raw data? 

Only Ooak. The raw archive never leaves Ooak's own pipeline and is never handed to another company. Our QA team works on the anonymized output, not the raw files. The buyer sees the code-named dataset and nothing else.

- 

### We are winding down. Who can actually sign? 

Whoever holds signing authority for the entity at that moment: a director, a liquidator or an administrator. We have closed deals in each of those situations. We ask for the authority document up front so the contract is not challenged later. Payment goes to the entity, not to individuals.

- 

### What happens to my raw data afterwards? 

It is deleted within 30 days of delivery. That is a contract obligation, not a policy, and you receive written confirmation. Only the anonymized dataset is retained for the licence term.

- 

### How much can I earn, and how is the number set? 

Deals to date range from $20K, the smallest closed, to several hundred thousand dollars. The number depends on years of history, headcount, the number of tools and how deeply they connect to each other. We never quote a number before counting. The estimator gives a range; the audit gives a figure.

- 

### What does this cost me in time? 

About one hour in total. One 15-minute intro call, one scoping pass by call or async, and either a native export from each tool or read-only access. No engineer is needed on your side. We do the extraction ourselves.

- 

### What kind of data are you looking for? 

We acquire the operational, strategy, and business data that defines how companies actually work. This includes documents and files (reports, contracts, presentations), project management data (Jira, Notion, Trello), CRM and sales data, financial records, product analytics, internal wikis, and code repositories. We are interested in the full context, not just isolated files, but the relationships between tools, teams, and workflows.

 - 

Next step

## Ready to find out what your data is worth?

We will walk you through the process, answer your questions, and give you an honest valuation. No commitment required.

 [Book a call](https://calendar.google.com/appointments/schedules/AcZssZ142jxQGEjB1PR9skAe17Nf7Pjf8G8EszSPcRhZIzmZXGCkn05REmmAmqPvMgWTYUMwbQNwinIF) [Email us directly](mailto:christian@ooakdata.com)