CTGT changes what a model does without touching its weights
Every team that has shipped an LLM feature knows the ritual that follows the first bad output.
The model invents a policy that doesn't exist, or refuses a harmless question, or picks up a bias nobody asked for. So the team rewrites the system prompt. When that stops working, they fine-tune, which takes weeks, costs real money, and produces a new model that has to be evaluated from scratch. When that still isn't enough, they bolt a second model on top of the first to judge its outputs, adding latency and another probabilistic component that can itself be wrong. Each layer of the stack is a patch over a system nobody actually understands.
CTGT starts from a different premise: if the unwanted behavior lives inside the model, fix it inside the model. The company, founded in mid-2024 by Cyril Gorlla and Trevor Tuttle and part of Y Combinator's Fall 2024 batch, builds what it calls the deterministic layer for frontier intelligence. Its platform locates the internal features that drive a specific behavior, whether that's hallucination, censorship, or a learned bias, and adjusts them directly at inference time, with no retraining and no guardrail model watching from the outside.
Investors moved quickly on the idea. In February 2025, Gradient, Google's early-stage AI fund, led an oversubscribed $7.2M seed round, joined by General Catalyst, Y Combinator, Liquid 2 Ventures, and Deepwater. The angels include Keras creator François Chollet, Y Combinator's Michael Seibel, Paul Graham, and Mark Cuban.
The patch stack is the problem
To see why CTGT's approach is interesting, look at what enterprises currently do when a model misbehaves.
Fine-tuning is the heavyweight option. It can genuinely shift behavior, but it's slow, expensive, and blunt. You don't tell the model what to stop doing; you show it thousands of examples and hope the gradient finds the lesson. Along the way it can forget things it used to do well, so every fine-tune triggers a fresh round of evaluation. For a bank or a hospital, that cycle can take a quarter.
Prompt engineering is the lightweight option, and it's brittle, as anyone who has maintained a 2,000-word system prompt can attest. Instructions interact unpredictably. A clause added to stop one failure quietly causes another. And a prompt can only ask the model to behave; it can't make it.
The third option, wrapping the model in classifiers and judge models, turns one inference call into several. Latency goes up, cost goes up, and the watchdog is a black box policing another black box. None of these approaches tells you why the model hallucinated in the first place. They manage symptoms.
CTGT's pitch is that all of this becomes unnecessary once you can see what's happening inside the network. The company puts a number on the speed difference: its approach to customizing and deploying models is up to 500x faster than traditional training, because it skips the training loop entirely.
Editing behavior at the feature level
The core of CTGT's platform is a technique the company calls feature-level model customization.
Inside a large model, behaviors correspond to patterns of internal activity. Somewhere in the network are latent variables that activate when the model is about to refuse a question, or drift from its source material, or apply a bias it absorbed in training. CTGT's tooling identifies those variables for a given behavior, then modifies them while the model runs.
Three properties make this different from everything else on the market.
The weights never change. The adjustment happens at inference, like a filter applied to the model's internal activity rather than surgery on the model itself. That makes every change instantly reversible and means the base model's capabilities stay intact. There is no risk of a fine-tune degrading performance on tasks you didn't think to test.
The change is deterministic. A prompt instruction is a request the model may or may not honor. Suppressing the feature that drives a behavior removes the behavior at its source. For enterprises that need to make reliability commitments, the difference between "the model usually complies" and "the circuit that produced that failure is switched off" is the difference between a demo and a deployment.
The interpretability is built in, not bolted on. Because the method works by locating the features behind a behavior, the explanation of what changed comes with the fix. CTGT describes its interpretability as mathematically verifiable, with no supplemental monitoring models required. For a compliance officer who has to document why the system behaves the way it does, that audit trail is the product as much as the fix is.
The DeepSeek demonstration
CTGT's most vivid public demo arrived with DeepSeek-R1, the Chinese open-source reasoning model that impressed engineers and alarmed compliance teams in roughly equal measure.
DeepSeek's raw model refuses or deflects a large set of sensitive topics. CTGT applied its method to the censorship behavior and reported that the model went from answering 32 percent of sensitive questions to 96 percent, later reaching 100 percent in its tests, with no degradation on neutral benchmarks like reasoning, math, and coding. The team has since run the same process on other open-source models, including Llama.
The censorship numbers made the headlines, but the underlying point is bigger than any one model's politics. Open-source models are attractive to enterprises for cost and control, and every one of them ships with behaviors its publisher chose, some documented, some not. A tool that can find those behaviors and switch them off turns open-source models from take-it-or-leave-it artifacts into something an enterprise can actually inspect and shape. The same mechanism that removes censorship can suppress hallucination, and CTGT says its technology reduces hallucinations by 80 to 90 percent, enough to support what it describes as 99.9 percent reliability for enterprise deployments.
Who buys this
CTGT sells where the cost of a wrong answer is highest.
The company says its platform is deployed with Fortune 500 companies, including tier-1 financial institutions and global media conglomerates, and that it has worked with a Fortune 10 company on safe, on-device AI. Its stated focus is high-stakes deployments in finance, healthcare, and security, the industries where AI pilots most often stall at the proof-of-concept stage because nobody will sign off on a system that occasionally makes things up.
One named customer makes the value concrete. Ebrada Financial used CTGT to improve the factual accuracy of its frontline customer-service chatbots, and its founder credits the platform with eliminating most of the escalations to human agents that inaccurate answers used to generate. That is the shape of the ROI story across the customer base. A hallucination in production has a price: the support ticket it opens, the compliance review it triggers, the customer it loses.
The on-device work points at a second market. Running models locally, on phones or in cars or inside air-gapped environments, means shipping behavior you can't easily patch from a server. A method that bakes reliability into the inference layer, without ballooning the compute budget the way judge models do, fits that constraint unusually well.
A founder who started early
Cyril Gorlla has been working on one question, why neural networks do what they do, for most of his short career.
He got into machine learning research at UC San Diego, where he met co-founder Trevor Tuttle, and left his research at Stanford at 23 to start the company. CTGT is that research agenda turned into infrastructure. Interpretability work mostly lives in academic papers and safety-team blog posts; Gorlla's bet is that it is actually the missing operational layer of enterprise AI.
The seed round's composition says investors with very different vantage points buy the thesis. Gradient and General Catalyst know what enterprises are asking for. Chollet, who built one of the most widely used deep learning frameworks, can judge the technical approach itself. Both camps wrote checks.
Interpretability becomes a product category
For years, interpretability research has been AI's good intention: everyone agreed it mattered, and almost nobody's production stack included it. The field produced striking papers about features and circuits inside large models while the industry shipped black boxes wrapped in disclaimers.
CTGT is a bet that this work is ready to be sold rather than just published. If models are going to run bank chatbots and clinical workflows, someone has to be able to answer the question "why did it say that, and how do we make sure it never says it again" with something better than a new system prompt. Regulators are starting to ask that question in writing. Enterprises are learning to ask it before the pilot rather than after the incident.
The companies that can answer it won't necessarily be the ones training the biggest models. They may be the ones who understand the models best. CTGT raised $7.2M on that wager, and its early customer list suggests the market is coming around to the same view: before enterprises buy a bigger model, they want a steering wheel for the one they have.