Ocular AI pays doctors and linguists to train voice models

Ryan Bednar9 min read
Ocular AI pays doctors and linguists to train voice models

The data frontier models need was never posted online

Every large language model you have used was raised on roughly the same diet: the public internet, plus whatever licensed text its lab could buy. That worked for a decade because the internet is enormous and nobody had finished eating it. The meal is mostly over now. The labs have scraped what there is to scrape, and the capabilities they want next live in places a crawler cannot reach.

Think about what a model needs in order to hold a real conversation. Not to answer a written question, but to talk. It has to hear a person trail off and know whether they are finished. It has to interrupt politely, and survive being interrupted. It has to catch the hesitation in a patient's voice when they insist the pain is "fine." Almost none of that exists on the web in usable form. Podcast audio is edited and mixed down. Call-center recordings are locked inside companies and wrapped in privacy law. And the judgment of a working expert, the way a doctor narrows a diagnosis or a lawyer hears the actual problem inside a client's rambling story, mostly exists in people's heads.

Ocular AI is built to get it out of their heads. The company describes itself as an applied AI data research lab, and its one-line mission doubles as the category it wants to own: encoding human expertise into machines, starting with voice. Founded in 2024 by two Dartmouth-trained engineers out of Google and Microsoft, and backed by Y Combinator in its Winter 2024 batch, Ocular recruits credentialed experts, records what they know in structured sessions, and runs the results through purpose-built infrastructure until they come out the other side as training data a frontier lab can use.

The next decade is a data problem

Ocular's founding thesis, laid out in the company's manifesto, is easy to state: the last decade of AI was won by compute, and the next will be won by compute and the right data, scaled together.

That second clause is doing a lot of work. Compute kept getting cheaper and models kept getting bigger, and for years the data side of the equation took care of itself because the web kept growing. What changed is the kind of capability labs are now chasing. Reasoning through a hard medical case, negotiating in real time, speaking naturally with a stranger who has an unfamiliar accent: these are skills, not facts. The web documents what people know. It rarely captures how experts think, and it almost never captures how they sound while thinking.

Ocular's manifesto argues that frontier models need to learn from the full fabric of human expertise: voice, vision, reasoning, taste, judgment, and the subtle ways experts work through hard problems. If that data does not exist online, someone has to manufacture it, deliberately, with real experts, at a quality bar a frontier lab will accept. That manufacturing process is the company.

What expert voice data actually looks like

Voice is Ocular's first market, and the specifics show why "just record some conversations" does not cut it.

One of the company's core offerings is full-duplex conversational data: two speakers captured at 48 kHz on isolated channels, with the overlaps, backchannels, and natural disfluencies preserved rather than cleaned away. That preservation is the point. A model trained only on tidy, one-speaker-at-a-time audio learns a version of conversation that does not exist in the wild. Real people talk over each other, mutter "mm-hm" while the other person speaks, restart sentences, and leave thoughts hanging. A voice model that cannot handle any of that sounds like what it is, which is software waiting for its turn.

The catalog goes deeper than general chat. Ocular builds domain-specific speech datasets around task-anchored sessions, including medical intake, customer support, technical interviews, and emergency calls, each tagged by scenario and intent. It also sells prepared datasets directly, such as a multi-accent English corpus for speech recognition, aimed at the stubborn gap between how benchmarks say a model performs and how it performs on English as most of the world actually speaks it.

Raw audio is only half the product. Ocular transcribes, diarizes, and annotates the recordings, down to disfluencies and prosody, then shapes everything into formats that drop into a lab's training pipeline. The company says it is already doing this work with leading AI labs building voice models. That customer base is telling: the organizations with the most compute and the best researchers in the world are the ones buying data from an outside specialist, because collecting it well is a different discipline from training on it.

Experts, vetted and paid like experts

The first era of AI training data ran on anonymous crowdwork. That was the right tool for its moment. Drawing boxes around pedestrians or rating a chatbot reply as good or bad does not require a professional license, and the work was priced accordingly.

The data Ocular sells cannot be produced that way. You cannot crowdsource a realistic medical intake conversation from people who have never taken a patient history, and you cannot annotate a technical interview well without knowing which answers are actually good. So the company's first pillar is a curated expert network: doctors, lawyers, engineers, linguists, researchers, creatives, and native speakers, matched to the tasks where their specific expertise matters. Contributors are credentialed and vetted before they work, paid commensurate with their skill, and tracked through every contribution they make.

That last detail matters more than it might seem. Provenance and consent are built into the pipeline, so a lab buying an Ocular dataset knows who produced each piece of it and on what terms. As scrutiny grows over where training data comes from, a dataset with a clean chain of custody is worth more than one with a murky one, independent of its contents. Ocular is betting that expert data with receipts becomes the standard frontier labs are held to.

The foundry underneath

The second pillar is the machinery. Ocular calls it the Data Foundry: infrastructure that ingests, processes, structures, and evaluates expert contributions at scale, turning raw human effort into datasets with quality checks and evaluation built in.

The company earned this plumbing the long way. Michael Moyo and Louis Murerwa started out in workplace search, frustrated that the knowledge inside a company was scattered across dozens of SaaS tools, and that problem pulled them into the deeper one of making unstructured data usable at all. From there Ocular shipped Foundry, an annotation platform with dataset versioning and review workflows, and Bolt, a service that put vetted domain experts to work on hard labeling problems in fields like medical imaging and legal documents. In the summer of 2025 the company launched a multimodal data lakehouse for ingesting, cataloging, searching, and annotating video, image, and other unstructured data.

Each of those products reads, in hindsight, like a component of the current one. The annotation tooling, the expert marketplace, the storage and retrieval layer for messy multimodal data: stack them up and you have a foundry. What changed is the sharpness of the aim. Rather than selling tools for other companies' data problems, Ocular now points the whole apparatus at the customers with the deepest data problem of all, the labs training frontier models.

Two scholarship kids from Lusaka and Harare

Ocular's founding story would stand out even if the company were ordinary.

Moyo, from Zambia, and Murerwa, from Zimbabwe, landed at Dartmouth on full scholarships, where Moyo studied biomedical and computer engineering and Murerwa studied computer science. After graduating, Moyo went to Microsoft and Murerwa went to Google in New York, where he worked on the distributed systems underneath Google Cloud. They reunited in 2024 to start Ocular and go through Y Combinator, and African tech press covered the acceptance as a first: no Zimbabwean-founded startup had made it into the accelerator before.

It is hard not to notice that a company selling multi-accent English speech data was founded by two people who grew up speaking Englishes that speech recognition has historically served worst. The team knows from experience where the gaps in the data are, because they have personally been on the wrong side of them.

The company raised a $2M seed round in November 2024 from investors including Drive Capital, BDMI, Alumni Ventures, and Y Combinator, along with Orange Collective, the fund of YC alumni that backs breakout companies from the accelerator's ecosystem. The team is small, under ten people, which for a data company is less a constraint than a stance: the leverage is supposed to come from the network and the foundry, not from headcount.

Voice first, then everything else

The sequencing of Ocular's mission statement is deliberate. Encoding human expertise into machines is the destination. Starting with voice is the wedge, and the timing explains why.

Voice is becoming a primary way people use AI, and the major labs are moving voice-first, shipping models that listen and speak natively rather than bolting transcription onto a text model. Every one of those efforts needs exactly what Ocular produces: large volumes of real conversation, captured to spec, annotated by people who know what they are hearing. It is a seller's market for a scarce input, and the company is already in it.

But nothing about the expert network or the foundry is specific to audio. The same pipeline that records a doctor's intake conversation can capture how a radiologist reads a scan, how an underwriter prices a risk, or how a senior engineer reviews code. Whichever expertise the labs need encoded next, the machinery transfers.

The scraped web took AI a very long way, and that era paid off spectacularly. The rest of what people know is not lying around online waiting to be crawled. It will have to be recorded on purpose, from the people who have it, with their consent and at professional quality. Ocular AI is building the company that does the recording.

Related Posts

Strada does the insurance work that starts after the call

Insurance still runs on phone calls, but most of the work happens after someone hangs up: updating the policy system, filing the claim, sending the follow-up. Strada builds AI agents that handle the call and the paperwork behind it, for carriers, MGAs, and brokers.

9 min read