By now, everyone has probably had an AI experiment that turned out to be disappointing. An
generated email that sounded like everyone and no one. A summary that prioritized the wrong things. A “smart” suggestion based on last year’s product catalog. HubSpot measured this across its entire customer base and arrived at a conclusion that fits into a single sentence: poor context is worse than no AI at all. Teams that used AI with poor context booked up to 49% fewer customer meetings and closed up to 70% fewer tickets than teams that didn’t use AI at all. Teams with good context did the opposite: 264% more MQLs, 224% more closed deals.
I had a front-row seat this week at Partner Day in Boston when HubSpot presented its solution to that problem: Aviator, an agent harness deeply embedded within the Smart CRM. In this article, I’ll explain what it is, why it’s more than just a chatbot on your CRM, how it fits in with what the rest of the market is doing, and what I, as an AI specialist, think of it—including some caveats.
What exactly is an agent harness?
A language model on its own doesn’t do anything. It answers a single question and stops. To get it to do work—such as researching a prospect, updating a deal, or setting up a campaign—you need a loop: gather context, let the model reason, check if the task is complete, and if not, go through another round. That mechanism is called an agent harness.
Anyone who’s built agents themselves knows that this loop is the hard part—not the model. Each round costs tokens—and thus money. Incorrect or excessive context makes the answer worse and more expensive. And if the harness is disconnected from your data, it knows nothing about your permissions, your approval workflows, or who switched roles last week.
What HubSpot Does Differently
HubSpot didn’t just add AI to its CRM—it rebuilt it from the ground up over the past year with a harness built right in. Three choices stand out.
-
Model-agnostic, yet decisive. HubSpot calls it “the Switzerland approach”: no single model is the best at everything, and that changes almost weekly. That’s why Aviator selects the model that performs best for each task, based on evaluations that HubSpot continuously runs using its own go-to-market expertise. If another model performs better for a task, Aviator switches to it. You won’t notice a thing, except that the output gets better.
-
The right context at the right time. Aviator works in conjunction with the Growth Context Graph: a layer that derives insights from all your structured and unstructured data and organizes them across three dimensions. Business context (brand, tone of voice, products, positioning), team context (roles, goals, approvals, workflow), and customer context (ideal customer profile, communication history, contracts, revenue). For each task, Aviator injects only the relevant portion. Not everything, because noise consumes tokens and lowers quality.
-
Built into the trust layer. Every agent running on Aviator is subject to the same security (role-based access), governance (permissions, approval workflows), observability (what each agent did, and why), and quality control. That’s the difference between this and an agent a colleague threw together on a Friday afternoon: that one works only until the person who built it leaves. HubSpot told that story themselves on stage, about their own team.
How this fits into what the market is doing
What HubSpot is building here doesn’t stand alone. It’s the CRM interpretation of a
shift that the entire AI industry has undergone over the past year.
From prompt engineering to context engineering. In 2023 and 2024, the focus was on the
right prompt. Since 2025, the industry has been talking about context engineering: the art of determining what information a model receives at any given moment. The labs themselves are saying it out loud. Anthropic published guidelines for building effective agents, concluding that simple, well-fed loops yield better results than complex frameworks. Aviator is exactly such a loop, with the context layer as its core component.
Models are becoming interchangeable, but the harness is not. Anthropic’s frontier models,
OpenAI, and Google’s frontier models catch up with each other every few months. Anyone who ties their product to a single model loses out. That’s why everyone serious about agents is building a model-agnostic layer with their own evaluations: Salesforce is doing it with Agentforce, Microsoft with Copilot Studio, and the major clients I speak with are building exactly the same thing internally. HubSpot’s decision to make Aviator model-agnostic is therefore not an innovation—it’s the industry standard. The difference lies in the evaluations: HubSpot tests for go-to-market tasks, something a generic layer cannot do.
Open protocols are winning out. Over a year ago, HubSpot was the first CRM to adopt MCP (Anthropic’s open protocol that enables AI assistants to communicate with tools), and it’s now the most widely used CRM connector in both Claude and ChatGPT. That’s important for you: it means your HubSpot data and agents work not only within HubSpot, but also through the AI tools your team already uses. The market is moving away from walled gardens; HubSpot is moving with it.
Agent sprawl is the next problem. Now that anyone can build agents, companies are seeing uncontrolled proliferation that no one can keep track of. I see it in virtually every organization I visit: dozens of agents and automations, scattered across tools, with no owner and no logging. HubSpot’s Agent Hub is a response to this within the CRM domain. It doesn’t solve the problem company-wide, but it’s one of the first platforms to even acknowledge the problem.
My assessment as an AI specialist
I’m generally positive, with three caveats.
What’s strong: HubSpot has understood that the value lies not in the model itself but in the combination of the harness, context, and governance. That’s the right architecture, and it aligns with what I’ve observed when building agents in other sectors: the model accounts for 20% of the work, while context and guardrails make up 80%. The fact that HubSpot includes the trust layer by default saves customers months of development work.
Note 1: The figures are HubSpot’s own data. The good/bad context percentages come from HubSpot’s usage data and represent correlations, not a controlled experiment. Companies with good context are likely also companies that have their affairs better in order to begin with. The general direction is correct, but I wouldn’t include the exact percentages in a business case.
Note 2: Convenience comes at a price. Aviator selects models, injects context, and
processes the data without you seeing a thing. That’s convenient, but it also means you become dependent on HubSpot’s choices. With every implementation, ask about observability: can you see what context an agent used and why? HubSpot says you can. Test it.
Note 3: Context is work, not a setting. This is the crux of the matter. Aviator can only perform well if the context is right: which sources are up to date, which data is outdated, who’s authorized to approve what, what is your brand’s voice, and who exactly is your ideal customer? HubSpot now provides a space for this (Context Home, with a score and recommendations), but filling it out and maintaining it requires human effort and domain expertise.
What this means for your organization
The order is crucial. The numbers show that turning on AI before your context is in order actively worsens your results. That’s why Six & Flow will now start every AI implementation on HubSpot with context: first the data foundation and context configuration, then the agents. Not because it’s more exciting, but because it’s the only sequence that supports the data, and because we’ve seen exactly the same pattern in other industries.
Want to know how your context is shaping up? That’s exactly the conversation I’d love to have with you. Schedule an appointment here!
