Context engineering now matters more than which AI model you choose. That is the core message from Jeff Dean, Google’s former Chief Scientist, in a recent Y Combinator interview with Diana Hu. The model itself is just one component. What determines real outcomes is the system built around it: how the model retrieves information, what tools it can use, and how AI agents coordinate to solve problems.
Why Is Model Choice Becoming Less Important?
AI progress used to mean training bigger models on more data. Jeff Dean told Y Combinator interviewer Diana Hu that this framing is increasingly outdated. The gains are now coming from what surrounds the model: retrieval (pulling in relevant, up-to-date information on demand), tools (actions the model can trigger), memory (what it retains from past interactions), and agent coordination. Dean said these elements “might get consolidated into what people call context engineering.”
His exact framing: “AI progress used to mean just better models. You had more data, train bigger models with bigger parameters. But increasingly in the last years or so, it’s everything around the model.”
Dean’s view is that the model people choose to use is just one piece of the system. What matters more, he said, is equipping the model with the right tools and the right access to relevant information, rather than expecting the model itself to carry all the weight.
What Does “Context Engineering” Actually Mean?
Context engineering is the practice of designing what information and tools an AI model sees and uses when working on a specific task, rather than relying solely on what it absorbed during training. Dean contrasted the two directly:
“The nice thing about that is that information is really clear to the model, unlike the training data the model is trained on where it’s all kind of like trillions of tokens stirred together into a soup of hundreds of billions or trillions of parameters. But it’s all less clear than the actual context that the model sees directly for this particular problem or use case.”
In plain terms: training data is vast but diffuse. Direct context is narrow but precise. When a model can pull in exactly the right documents or data at the moment it needs them, it performs better than a model relying purely on recall from training.
Why Is Multi-Agent Orchestration Becoming Central?
Dean extended the context engineering idea into multi-agent orchestration (coordinating multiple AI agents, each with specific tools and roles, to solve complex tasks together). He described the capability as understanding “what tools are available, which ones are going to help the model solve this next phase of the problem, how to decompose the problem into a sequence of tool calls, maybe trying multiple approaches to solve the problem and seeing which ones work.”
His conclusion: this kind of orchestration is where AI capability is heading, and that it represents “super exciting times.”
What Are Jeff Dean’s Practical Tips for Better Context Engineering?
One point Dean stressed: context engineering, unlike model training, does not require GPU infrastructure or massive compute budgets. Anyone with API access, such as to Gemini, can practice it. His tips focused on a learn-from-failure, iterate-forward approach.
Rather than trying to adjust model parameters (which most practitioners cannot do), Dean recommends writing better guidelines and building better skills, meaning structured instructions that tell the model how to use specific tools for specific types of problems. He described this as a “self-improving setup” that gets better over time as you identify and address failure cases.
“Often you can actually make the model work better and succeed at that kind of problem by not just adjusting the model parameters, which is hard to do from the outside, but from creating better guidelines for the model, writing skills for the model to know how to use different tools that would be incredibly useful for solving this particular class of problem.”
A starting checklist based on Dean’s framing, for anyone building or auditing an AI-powered workflow:
- Pick one specific task where your current AI setup is underperforming, not a vague category.
- Log concrete failure cases, not general impressions.
- Determine whether the model lacked the right information at the right moment (a retrieval gap) or reasoned incorrectly from information it did have (a reasoning gap).
- Write clearer guidelines for the task: specify what information the model should pull and when.
- Define which tools the model should use at each stage of the task, explicitly.
- Re-run the same failure case with the updated setup and compare outputs before drawing broader conclusions.
- Repeat across several failure types before deciding the model itself is the problem.
What This Means for AI-Search Visibility
Here is the simplest way to understand the business implication: a brand that appears in an AI’s training data is fighting last year’s battle. A brand that appears in the context an AI actively retrieves is competing in today’s one.
Dean’s description of training data as “trillions of tokens stirred into a soup” is a useful mental model for why training-data presence alone is not a reliable visibility strategy. It is diffuse. Retrieval is specific. If an AI assistant is helping a user research a product or service, it may retrieve current web content or structured data in real time. A brand’s presence in that retrieved layer matters more than whether it appeared somewhere in a training corpus.
One implication the source does not explicitly draw out, but is worth naming: Dean’s framing also applies to how content is structured. Structured, clearly attributed, and self-contained content gives retrieval systems more to work with than broad or vague content. A page that makes a specific, sourced claim in a single paragraph is easier for a retrieval system to pull and surface accurately than a page that buries the same claim inside 2,000 words of background.
Hingewise’s view: the most actionable takeaway from Dean’s framing is not “which model should I use” but “what does my content look like from the perspective of a retrieval system.” That question is closer to classic information architecture than it is to prompt engineering, and it is one that brands and publishers can act on now, regardless of which AI platforms they are targeting. How retrieval architectures evolve across different AI search products over the next 12 to 18 months will determine how much that action translates into measurable visibility gains.
Source: Search Engine Journal / Original interview: Diana Hu, Y Combinator
Lam Nguyen · Hingewise
