A customer cancels their subscription.
Your growth stack has work to do: record the cancellation, stop inappropriate campaigns, preserve any remaining paid access and update reporting.
An AI agent could read the account history, interpret the cancellation message and recommend an intervention.
But should it decide when access ends? Change the customer’s lifecycle state? Offer a discount? Send an email?
Those are different permissions. Treating them as one capability is how an impressive demonstration becomes an unreliable operating system.
The useful question in AI agents vs automation is which parts of the process need interpretation, which need dependable execution, and who has authority to change the outcome.
For founders, growth leaders and marketing operations teams, this is becoming an architectural decision: what should AI be allowed to observe, decide and do inside the business?
When should you use an AI agent instead of automation?
Use an AI agent when completing a task requires choosing the next step from changing context, gathering additional evidence or adapting tool use—and when that flexibility creates measurable value within enforceable limits.
Use conventional automation for known rules. Use a bounded AI step when interpretation is needed but the process is already known. Keep consequential policy decisions under human ownership.
This qualifies the familiar argument that the future growth stack will be hybrid.
Deterministic software can enforce reliability controls; it does not guarantee reliability. Agents provide probabilistic interpretation and planning; they do not acquire accountability. Humans remain accountable for delegated actions even when they do not approve every execution.
The practical principle is:
Use rules for established policy. Use AI for useful interpretation. Use agents for useful adaptation. Authorize actions separately.
Research note: this investigation draws on technical documentation, engineering accounts and independent research reviewed in September 2026. The decision frameworks are HGW analysis. The lead-qualification comparison is a proposed architecture, not a completed HGW test.
Why deterministic automation is not obsolete
For a purchase event, the required process may already be explicit:
- Validate the event.
- Resolve the customer’s identity.
- Record payment status.
- Remove prospect-campaign eligibility.
- Enter the appropriate onboarding journey.
- Send the agreed event to analytics.
An agent does not improve that sequence merely by choosing to perform it.
Predefined logic is easier to test against a policy, constrain for latency and inspect after failure. Adding model calls introduces another dependency and another source of variation.
This matters for suppression lists, entitlement changes, event transformations, synchronization and exact campaign eligibility.
Repeatable does not mean correct
A perfectly repeatable workflow can consistently apply a bad rule. It can also receive duplicated webhooks, race against another update or lose an API response.
Reliability requires correct policy, sound data and failure handling, whatever chooses the next action.
Even the cancellation example needs precision.
Cancellation requested, renewal disabled, subscription expired and access revoked are not necessarily the same event.
Someone who cancels renewal today may remain a paying customer until their current period ends. Removing them from every active-customer segment—or immediately revoking access—could be wrong.
The business must define those transitions before automating them. That is the foundation explored in How to Design Customer Lifecycle States.
Intelligence has an operating cost
Where the business logic is known, an additional interpretation layer must justify its existence.
Consent checks should follow explicit policy. A verified purchase should stop an inappropriate checkout-recovery sequence. A scheduled subscription expiry should follow the authoritative billing state.
Do not introduce judgment where judgment is not required.
The qualification is that AI may help write, test or maintain these rules without participating in every live execution.
Using AI to build an automation and using AI to decide what that automation does at runtime are different architectural choices.
AI agents vs workflows: define who controls the next step
“AI agent” has become an elastic product label. An operational definition is more useful than a marketing category.
Anthropic distinguishes workflows with predefined code paths from agents whose models dynamically direct their processes and tools. OpenAI similarly defines agents around model-controlled workflow execution, rather than treating every LLM application as an agent. These are useful definitions, not a universal industry standard. Anthropic’s architectural guidance, OpenAI’s practical guide.
| System | Who chooses the execution path? | Growth example |
|---|---|---|
| Deterministic automation | Predefined code or workflow rules | Route a verified enterprise lead by territory. |
| AI-assisted workflow | A fixed process containing bounded model tasks | Classify free-text intent, then apply routing rules. |
| AI agent | A model selects subsequent steps within permitted boundaries | Investigate an account, choose relevant sources, resolve gaps and propose routing. |
| Human-in-the-loop system | Software prepares work; a person approves defined actions | Review an exact retention offer before sending. |
| Human-led decision | A person determines the decision; software supplies evidence | Change pricing policy or approve a strategic partnership. |
These categories overlap.
A deterministic workflow can call an agent. An agent can call a deterministic workflow. Both can require human approval.
“Agentic automation” describes that combination, but tells you little about its actual authority. Ask what the model can choose and which changes it can cause.
This is not a maturity ladder
The sequence Deterministic Automation → AI-Assisted Workflow → AI Agent → Human Decision is useful as a selection menu.
It is misleading as a maturity ladder.
Human decisions are not “more automated.” Greater ambiguity does not automatically justify greater autonomy. A sophisticated business may deliberately keep most customer-state changes under conventional automation.
Nor does every prediction require a generative agent.
Conventional statistical models can score churn, rank offers or detect anomalies. The relevant comparison may be rules versus a trained classifier versus one LLM call, before agents enter the discussion.
Observe, decide, act: three separate permissions
An agent’s access should be specified in three parts.
Observe: what may it read?
Potential sources include CRM records, lifecycle state, product events, support conversations, email engagement, payment history, company enrichment and experiment results.
Each has different sensitivity and freshness requirements.
An account summary rarely needs raw payment details or every employee’s correspondence. A lead-research agent may need a company website without needing access to the entire customer database.
Decide: what may it infer or recommend?
Classifying buying intent, selecting an investigation path and recommending an offer are different from establishing a customer’s contractual status.
The agent might infer that an account needs attention. That does not make its assessment the authoritative customer state.
Act: what may it change?
Creating an internal task, modifying a segment, triggering a campaign, changing access and issuing a discount have different consequences.
They should not arrive as one bundled permission.
| Capability | Permission granted | Permission still withheld |
|---|---|---|
| Observe | Read approved account context | Read unrelated accounts or sensitive fields |
| Decide | Propose intent and routing with supporting evidence | Redefine eligibility or commercial policy |
| Act | Write an assessment to approved fields | Change consent, billing state or access |
Permission to observe does not imply authority to decide. Permission to decide does not imply authority to act.
Reading can itself create risk through disclosure. “Read-only” is not equivalent to harmless.
Give every action a contract
HGW recommends defining an action contract for each executable tool:
- Who may invoke it?
- Which customers and fields may it affect?
- What evidence must exist?
- Which current-state checks must pass?
- What limits or approvals apply?
- How are duplicate actions prevented?
- How is the result verified?
A tool named update_customer is too broad for many growth tasks.
A tool that records a qualification assessment against a specific lead, without changing consent or ownership, gives the system a smaller and more testable responsibility.
Generating copy, recommending that it be sent and sending it are three separate operations. A well-written draft proves none of the conditions required for sending.
Which growth tasks belong where?
Most growth functions contain rules problems, interpretation problems and accountability problems simultaneously.
Allocate individual decisions, not entire departments, to AI.
| Growth area | Rules or conventional models | Bounded AI or agent contribution | Human-owned boundary |
|---|---|---|---|
| Acquisition | Budget caps, tracking, pacing alerts | Creative analysis; investigation of campaign changes across sources | Budget strategy, claims, major campaign changes |
| Lead capture and qualification | Validation, deduplication, territory routing | Intent classification; adaptive account research | Qualification policy, strategic exceptions |
| CRM and customer data | Authoritative updates, exact identity matching | Enrichment proposals, ambiguous duplicate review | Destructive merges, disputed identity |
| Lifecycle marketing | Consent, suppression, timing, eligibility | Draft personalization; recommend an approved journey | Sensitive messaging, policy and brand approval |
| Conversion | Checkout triggers, offer eligibility | Sales summaries; contextual assistance | Discount limits, commercial commitments |
| Onboarding and activation | Milestone tracking, access provisioning | Interpret obstacles; investigate missing activation | Exceptional access, service commitments |
| Analytics and experimentation | Metric computation, statistical tests | Query assistance, hypotheses, investigation | Causal conclusions, experiment decisions |
| Retention and expansion | Renewal dates, validated risk models | Synthesize feedback; propose interventions | Substantial refunds, termination, account strategy |
| Growth operations | Synchronization, retries, scheduled reports | Summaries, documentation; bounded incident investigation | Production workflow and infrastructure changes |
This is a starting allocation, not evidence that every AI application listed is economically justified.
Where AI creates useful leverage
For qualitative feedback, a fixed model call may replace dozens of brittle keyword rules.
Google’s architecture guidance explicitly identifies summarization, translation and feedback classification as tasks that may not require agentic infrastructure. Google Cloud design guidance.
An agent becomes more plausible when the investigation itself changes.
A weak activation signal prompts a product query. That reveals a failed integration. The agent then checks support history before proposing an intervention.
The value lies in adapting the investigation to the evidence.
Where interpretation must stop short of authority
In experimentation, a plausible explanation is not causal evidence.
Let software compute the predefined analysis and AI explore interpretations. Keep changes to success metrics, stopping rules and rollout decisions explicit.
Otherwise, an agent can turn uncertain measurement into a confident narrative—and then optimize against it.
This extends the measurement discipline in Why Your Growth Analytics Cannot Be Trusted.
A practical framework for deciding when to use AI agents
Start with the cheapest credible baseline, then apply gates.
Do not average away a serious risk with a high “AI suitability” score.
| Dimension | Question | Architectural implication |
|---|---|---|
| Predictability | Are valid outcomes and steps known? | Prefer predefined logic. |
| Variability | Are inputs novel or unstructured? | Test bounded AI against rules or conventional models. |
| Interpretation | Is one classification enough, or must evidence gathering adapt? | Fixed inference suggests a workflow; adaptive investigation may justify an agent. |
| Reversibility | Can the customer impact actually be undone? | A reversible database edit may trigger an irreversible message. |
| Error consequence | What is the worst credible mistake? | Set action limits and approval requirements first. |
| Frequency | Is there enough volume to repay build effort? | Low volume may favour human-led work. |
| Observability | Can evidence, actions and outcomes be inspected? | Reduce autonomy if decisions cannot be audited. |
| Data quality | Are identity, freshness and ownership dependable? | Missing authority is a data problem, not an agent opportunity. |
| Permissions | Can access be scoped to this job? | Broad credentials can make deployment unacceptable. |
| Economics | Does interpretation improve total operating value? | Compare accepted outcomes, review and recovery costs. |
Apply this in order:
- Define the outcome and the governing policy.
- Test conventional execution.
- Add bounded interpretation if necessary.
- Add adaptive tool choice only when it improves results.
- Set autonomy according to consequence and evidence.
Poor data, inadequate permissions or unmanageable error consequences can stop deployment regardless of potential productivity gains.
“Use agents when interpretation is required” is therefore too broad.
Interpretation justifies testing AI. Adaptive execution justifies testing an agent. Neither automatically justifies permission to act.
The hybrid growth architecture: AI is an optional branch
A diagram that puts an agent between every customer event and every business system gives AI too much architectural importance.
The default path should remain available without it:
Validate the event → Resolve identity → Check policy → Execute the known workflow.
Send ambiguous cases to bounded inference or a restricted investigation loop. Return a structured proposal to the execution layer.
Routine work should not acquire a model dependency simply because another part of the process benefits from AI.
How Make, n8n and Zapier fit
This design is consistent with available tooling.
Make documents agents using modules, scenarios and MCP tools, with structured inputs and outputs. n8n documents approval gates on individual agent tools. Zapier documents both instruction-based approval and a separate Human in the Loop workflow step.
These are documented capabilities, not comparative reliability results. Make documentation, n8n documentation, Zapier documentation.
One implementation detail matters enormously:
An approval step after an agent cannot retrospectively protect changes the agent already made.
Withhold the write tool, or gate its execution before the side effect.
What orchestration and MCP actually contribute
APIs perform operations. Webhooks report events. Workflow engines coordinate execution.
MCP standardizes connections between model applications and external context or tools. It does not reconcile customer identity or establish your commercial policy. Its security guidance still requires careful authorization design. MCP specification, MCP security guidance.
Not every business needs a separate orchestration product. These responsibilities can live in application code or an existing platform.
Microsoft’s orchestration guidance illustrates multiple coordination patterns; it does not make a multi-agent system a prerequisite. Microsoft architecture guidance.
Human oversight and observability belong across this architecture. They are not a final box after customer impact.
Make vs agent vs hybrid: a proposed lead-qualification comparison
“Make versus an agent” is a false product boundary. Make can host conventional workflows and agentic steps.
The meaningful comparison is between execution designs.
Consider a form collecting:
- Name and email.
- Company and website.
- Job title.
- A free-text description of what the prospect needs.
The objective is to route legitimate inquiries promptly without losing ambiguous opportunities or inventing qualification evidence.
Version A: deterministic qualification
Validate the submission, resolve duplicates, call a fixed enrichment source, score defined attributes, update approved CRM fields and route by policy.
Missing information goes to a review queue.
The process is explicit. Its weakness is whatever the rules fail to capture.
Version B: agentic qualification
Give a restricted agent approved research tools.
It chooses what to investigate, follows relevant evidence, assesses fit and selects a routing proposal. In the proposed Lab environment, it writes to a sandbox CRM.
It has a fixed run budget and cannot contact prospects.
The potential advantage is adaptive investigation. The potential weakness is variable evidence quality and execution.
Version C: hybrid qualification
Deterministic infrastructure gathers and validates context.
One bounded AI call interprets the free text. A restricted agent investigates only unresolved cases. Policy gates constrain routing; uncertain or strategically important cases require review.
The workflow executes approved updates.
The hybrid need not include that agent branch if bounded interpretation is sufficient.
| Dimension | Deterministic | Agentic | Hybrid |
|---|---|---|---|
| Reliability | Stable rule application; policy gaps remain | Variable evidence and execution paths | Controlled writes; inference still needs evaluation |
| Flexibility | Limited to encoded cases | Can adapt investigation | Adapts selected exceptions |
| Build complexity | Rules and integrations | Prompt, tools, evaluation, runtime | More boundaries to implement |
| Maintenance | Rule and schema changes | Model, prompt, source and tool changes | Both, but smaller AI scope |
| Cost | Usually fewer variable calls | Depends on investigation length | AI spend concentrated on ambiguity |
| Latency | Usually most predictable | Long-tail research delays | Fast default path; slower exceptions |
| Transparency | Explicit scoring logic | Evidence and tool history needed | Separates assessment from execution |
| Scale | Mainly integration throughput | Also inference budgets and variable runs | Review queue may become a bottleneck |
| Failure modes | Brittle rules, missing enrichment | Unsupported inferences, wrong tools | Gate errors and poor escalation design |
| Edge cases | Escalates unless encoded | May resolve or confidently mishandle | Resolves within limits, otherwise escalates |
These are architectural expectations, not measured rankings.
Suppose a small agency writes: “We manage onboarding for 40 client companies.”
A company-size rule could underrate it. AI may recognize partner potential.
That does not authorize marking it “qualified enterprise buyer.” Store a partner-interest assessment with evidence and route it through an approved exception path.
Compare all three designs against the same historical cases and time-appropriate evidence. Measure legitimate leads missed, incorrect priority routing, unsupported claims, review effort, duplicate writes, cost and latency.
A system that routes everything to humans may look safe while saving no work.
What happens when the agent gets it wrong?
Agent failures combine familiar distributed-system problems with model-specific ones.
| Failure | What changes with AI? | Necessary response |
|---|---|---|
| Hallucinated facts and inconsistent decisions | The model may invent evidence or vary classifications | Require source references, unknown values and evaluated abstention. |
| Wrong tool or malformed action | The model selects the operation or arguments | Use narrow tools, schemas and server-side policy checks. |
| Stale or incomplete context | An ordinary data problem gains persuasive interpretation | Carry source timestamps; recheck authoritative state before writing. |
| Prompt injection | Customer text or retrieved pages may redirect behaviour | Treat external content as untrusted; restrict tools and data exposure. |
| Excessive permissions | Bad reasoning can cause broader damage | Enforce least privilege outside the prompt. |
| Loops and cost escalation | The model may keep investigating or retrying | Cap steps, time, tool calls and spend; define stopping conditions. |
| API failure and duplicated actions | Existing retry problems persist | Use bounded retries, idempotency and reconciliation. |
| Irreproducible decisions | Exact reruns may change | Preserve evidence and versions; test repeatability statistically. |
Valid JSON is not a valid business decision. A correctly formatted discount can still breach policy.
Deterministic guards can bound permitted actions, but cannot prove every recommendation is sensible.
Prompt injection turns evidence into an attack surface
An agent researching websites or reading inbound messages encounters content it should treat as evidence.
An attacker can place instructions inside that content.
Anthropic’s browser-agent research explicitly acknowledges residual risk despite improved defences. That supports containment, not a universal failure rate for all business agents. Anthropic prompt-injection research.
A prompt saying “ignore malicious instructions” is not an adequate substitute for limited data access and restricted tools.
A timeout does not prove the action failed
If a CRM write succeeds but the response times out, repeating the call can duplicate work.
AWS’s engineering guidance explains using explicit request identifiers to make retries safe. AWS Builders’ Library.
Carry an idempotency key through execution, verify destination state and track uncertain outcomes.
An agent’s statement that it completed a task is not a receipt from the destination system.
Recheck eligibility immediately before consequential writes. The customer may have purchased, unsubscribed or changed account owner while the agent was investigating.
This extends Marketing Automation Architecture: What Should Actually Be Automated?: entry conditions do not guarantee continued eligibility.
Human approval is an operating function
Some decisions deserve human ownership even if models become much better:
- Changing pricing.
- Issuing substantial refunds.
- Handling sensitive communications.
- Terminating customers.
- Approving strategic partnerships.
- Changing production infrastructure.
The reason is authority over acceptable trade-offs, not simply model weakness.
Approval works only when the reviewer sees the exact proposed action, affected customer, evidence, uncertainty and consequences.
Record who approved what. If the payload changes or approval expires, require a new decision.
A generic “looks good” should not unlock unrestricted later actions.
Human-in-the-loop can fail too
Queues grow. Reviewers rubber-stamp. Polished explanations conceal weak evidence.
Measure review time, overrides, missed errors and queue delay. Escalate material uncertainty rather than requiring approval for everything.
If the queue cannot be staffed, reduce scope instead of silently relaxing the gate.
Grant autonomy gradually
HGW’s graduated-autonomy model is:
Observe → Recommend → Prepare → Act with Approval → Act Within Limits → Autonomous Action.
The final stage means no routine approval within a defined mandate, not unlimited authority.
Progress should depend on demonstrated performance for a particular task and population. Roll back after drift, incidents or material system changes.
A self-reported “95% confidence” is not an escalation policy.
Test how observable signals—including missing evidence, source conflict and unfamiliar inputs—predict actual errors. Sample approved and unreviewed outcomes so evaluation does not cover only cases already flagged as difficult.
Agents need a customer model, not five contradictory profiles
An agent reading CRM, billing, analytics, lifecycle and support systems may receive five incompatible accounts of the same customer.
More context does not resolve authority.
Define which system owns each fact, how identities connect, how events are timestamped and how corrections propagate.
That is the foundation of What Should Be the Source of Truth for Customer Data? and How Customer Data Should Move Through Your Growth Stack.
Separate facts from assessments
“Subscription paid through 30 September” should come from billing.
“Likely to churn” should be stored as an assessment with evidence, creation time, expiry and model version.
Neither should silently overwrite the other.
Likewise, preserve the distinction between an observed product event and an inference about intent. A support conversation can inform an intervention without becoming an authoritative instruction to change access.
Agent memory is not automatically a source of truth. Old summaries require refresh, and one customer’s context must not leak into another’s run.
Adding AI to fragmented infrastructure can amplify errors by making contradictory data sound coherent.
Observability: how do you know what the agent actually did?
A chat interface is not an operational record.
Growth infrastructure needs a trace connecting the original event to the verified outcome.
Record:
- Event and customer identifiers.
- Context sources, timestamps and versions.
- Model and prompt versions.
- Tool calls and arguments.
- Policy checks and approvals.
- Errors, retries and uncertain results.
- Cost and latency.
- Confirmed destination changes.
OpenAI’s Agents SDK documents tracing for generations and tool activity, including controls over sensitive trace data. Tracing capability still needs configuration appropriate to customer information. OpenAI tracing documentation.
Decision summaries can help inspection, but generated explanations are not guaranteed faithful accounts of internal reasoning.
Audit evidence, execution and outcomes. Do not treat a persuasive explanation as proof.
Economics: compare cost per acceptable outcome
Token price is only one term in the operating model.
Total cost = implementation + maintenance + infrastructure + model/tool usage + human review + expected error and recovery cost.
Divide by acceptable completed outcomes, not attempted runs.
Track serious incidents separately. Averaging can hide unacceptable consequences.
An illustrative comparison
Assume 10,000 monthly cases, a deterministic execution cost of $0.002 and an agent cost of $0.08.
Execution totals are $20 and $800. If both deliver the same result, the agent adds $780 without corresponding value.
These are illustrative assumptions, not vendor prices.
Now suppose interpretation saves two minutes on 2,000 cases. That releases about 66.7 hours, worth $2,000 at an assumed $30 hourly labour cost.
If 20% of those cases need three minutes of review, that consumes 20 hours, worth $600.
The remaining $1,400 of labour capacity must cover incremental execution, maintenance and error costs before the design creates net value.
Time released is not automatically cash saved. Its value depends on whether the business redeploys that capacity, avoids hiring or reduces paid work.
A complex rule engine can also be expensive to maintain. Deterministic software does not always win economically.
Measure both paths. High volume can justify a cheaper specialist model or conventional classifier. Rare, complex inquiries may favour a person using AI interactively.
What is viable now, and what improves next?
Public evidence supports scoped deployment more strongly than unrestricted autonomy.
Measuring Agents in Production reports a survey of 306 practitioners and 20 case studies. Its production findings favour constrained execution and identify reliability as the leading development challenge.
Recruitment favoured deployed systems; results are descriptive, self-reported and not a randomized comparison of growth architectures. Production study.
Independent benchmark research also distinguishes average task success from consistency, robustness, predictability and error severity.
Better capability scores do not establish reliable customer operations. These benchmarks are not direct tests of marketing stacks. Agent reliability research.
Three categories are useful.
Established implementation options: bounded classification, drafting, summaries, tool calls and approval gates. They still require local evaluation; documented availability is not a reliability guarantee.
Increasingly viable, conditional on evaluation: bounded account research, operational investigations and contextual recommendations with verified tools and controlled writes.
Experimental as a general growth operating model: open-ended optimization across budgets, offers, customer access and production workflows without routine oversight. This research does not establish dependable end-to-end operation of that model.
Infrastructure is improving too.
Anthropic’s 2026 engineering account separates the reasoning loop, execution environment and durable session history. It also describes removing obsolete scaffolding as model behaviour improves.
This is an implementation account, not proof of autonomous growth performance. Managed Agents architecture.
HGW expects better models, context management, tool interfaces and evaluation to expand the range of economically useful delegated work.
That is a forecast. It does not eliminate authorization, data ownership or verification.
What should actually run your growth stack?
Start with one recurring task where people spend measurable time interpreting ambiguity.
Write its action contract. Establish a conventional baseline. Test bounded AI before adding adaptive tool use.
Deploy in shadow mode, then grant limited authority only when evidence supports it.
Before adding another layer, audit the existing growth stack. Better identity, clearer lifecycle definitions and simpler workflows may solve the problem already.
Agents are meaningfully different when models control the execution path. They remain software systems built from workflows, tools and state.
The label matters less than the authority delegated and the evidence that delegation works.
Use the least complex system capable of reliably performing the job.
Let rules execute established policy, AI interpret where it adds value, agents adapt where necessary, and people own consequential trade-offs.
As with a business that has outgrown its tech stack, sophistication is not the objective.
The best stack is the simplest coherent architecture capable of supporting the next stage of growth.


Leave a Reply