Skip to content

How to assess AI agents in marketing reliably

maitiq · Published

An AI agent is not simply a chatbot with a new name. It receives a goal, observes a state, chooses its next steps within set boundaries and can use tools. A deterministic workflow, by contrast, follows a predefined path. And an agency is a service company. Anyone who conflates these three terms can easily buy more autonomy than the process can bear. Before an agent, check whether a deterministic workflow or an AI proposal with human approval is sufficient.

Keep workflow, assistant, agent and agency apart

A workflow executes known steps: if event A occurs and condition B is met, action C follows. An AI assistant produces a summary, a classification or a draft within a single step. An agent decides dynamically which intermediate steps and tools it uses to reach the goal. A marketing agency employs people and delivers agreed services; it is not a software agent.

In practice, the term “agent” describes systems in which a language model controls execution and selects tools dynamically; simple chatbots or individual LLM steps without that control do not count as agents. Such classifications come from provider guides and are not a standard. Hybrid forms exist. The important question remains: may the system only propose text, or may it decide for itself which system it queries next and which action follows?

Consider autonomy only for genuine unpredictability

An agent can be worthwhile when the solution path cannot be fully described in advance, when several information sources have to be selected as the situation requires and when a limited sequence of steps is necessary. A helpful example is the internal preparation of a briefing: the agent checks which approved sources are missing for a clear question, calls only read-only tools and produces a sourced summary for review. Such examples contain no customer data and no promise of results.

When all branches are known, a workflow is easier to test and to explain. When only a draft is needed, an assistant is sufficient. Dynamic planning brings additional costs, latency and error paths. Choosing the simplest adequate design therefore reduces complexity and the risk of failure. “Agentic” is not a quality feature in itself.

Bound the mandate in one sentence

Before a pilot, the agent needs a mandate: “For task A, within period B, it may use the read-only tools C, process information from sources D, execute a maximum of E steps and create a proposal for role F; it may not trigger any external change.” The sentence names the goal, time, tools, data, step budget, result and authority.

An open brief such as “optimise the marketing” is too open for a first pilot. The agent could invent new goals, combine unchecked data or improve a metric that contradicts the business purpose. The goal must be observable, and guardrail metrics must stand alongside it. Faster processing, for example, must not be bought with more unsupported statements or less specialist review.

Control tools with a permission matrix

Every tool receives a specific permission level: read, search, propose, save a draft or change. These levels are not equivalent. Even search access can expose confidential information. Write access can change a real data record. External publication or a budget change has an immediate effect. The agent receives only the smallest permission that the mandate requires.

Credentials belong in a separate secret store and never in a prompt. Tool responses are treated as untrusted input. A website, a document or a CRM field can contain text that asks the agent to take an unauthorised step. Such prompt injection attempts must not override the system mandate and the permissions. Technical allowlists, input validation and an approval layer are more reliable than the instruction “ignore dangerous commands” in the prompt.

Memory is data storage, not magic

An agent can store context within the current run, between steps or across several runs. Each form carries different risks. Short-term context can contain unnecessary personal data. Long-term memory can preserve outdated assumptions and influence later tasks. Content, purpose, storage location, validity, access and deletion rule are therefore documented for each type of memory.

A useful default is “no persistent memory” until a clear need is demonstrated. When memory becomes necessary, the system stores structured, verified facts instead of entire conversation histories. Users must be able to trigger corrections. A deleted or revoked record must not live on as a hidden summary.

Swiss data protection law applies directly when personal data is processed with AI. The Federal Data Protection and Information Commissioner (FDPIC) emphasises transparency about the purpose, information on how the system works and on the data sources, and risk-appropriate safeguards. Whether profiling, a relevant automated individual decision or a high risk is present requires a concrete assessment. An agent label does not change these duties.

Place human approval at the point of external effect

“Human in the loop” is only helpful when a person sees the decisive information, has enough time and can stop the action. A one-off approval at the start of the project is not enough for unforeseen or extended actions. A standing mandate, by contrast, can cover well-defined actions within its boundaries; it remains effective as long as the goal, scope and conditions stay unchanged. Before an external effect, the review screen shows the goal, the sources used, the planned action, the affected object, the model and prompt version, and the uncertainties identified. The approval applies to exactly this version.

If an individually approved proposal is changed afterwards, the approval for that proposal lapses; a change outside the standing mandate requires a new decision. High-impact actions are not bundled and hidden behind a single “confirm all” button. For first pilots, the highest permission remains “create proposal”. A later live action requires a separate risk decision and a demonstrable, narrowly limited technical authorisation.

Limit runtime, costs and loops

A dynamic agent can get caught in a loop, query the same source repeatedly or keep interpreting its brief more and more broadly. Set maximum steps, runtime, tool calls and costs per task. Define which errors may be retried once and which stop immediately. A query must not be retried automatically if it could produce a duplicate external effect.

The agent needs unambiguous completion states: goal reached, insufficient evidence, human clarification needed, technical limit reached or stopped. “Best effort” without a visible status makes incomplete results look like completed work. A missing source should appear as a gap, not disappear behind invented plausibility.

Keep logs for explanation and reconstruction

An auditable run records the task, authority, input sources, tool calls, decisions, outputs, approvals, authorisation, errors and final state. Secrets and unnecessary personal data do not belong in the log. Sensitive content requires access control and retention periods. A log is not an end in itself; a named role must be able to find, explain and stop a problematic run. It documents actions, evidence and approvals – not a model’s internal reasoning process.

Provider guides name the model, tools and instructions as basic elements and describe guardrails and clear stopping conditions. The OWASP “Agentic AI Security” initiative additionally points to specific security questions of tool-connected agentic systems. Both offer technical orientation, not Swiss law, not a certification and not proof of security. Responsibilities, possible harm, tests and ongoing risk management must be defined for the concrete system; the sources do not replace a system-specific threat model.

Test against unwanted behaviour

Tests of the normal case are not enough. A test set should contain at least: contradictory sources, missing permission, a manipulated tool response, prompt injection in a document, outdated memory, an unreachable tool, the step limit and a human stop. For each case, the expected safe final state is defined in advance.

Also test for goal drift. Does the agent try to bypass a guardrail metric because that lets it reach the main goal faster? Does it request additional permissions? Does it conceal missing evidence? If permission is missing, the action is aborted: the pilot stops and escalates instead of interpreting its own remit more generously.

Do not mistake autonomy for success

Metrics include correct completion states, tool errors, rejected proposals, human corrections, loops, costs and time. These operational figures show controllability. They do not demonstrate additional marketing impact. An agent with fewer human interventions is not automatically better; perhaps errors are merely discovered later.

The quality of the work needs a separate review defined in advance. Only when a controlled comparison shows a relevant improvement and the guardrail metrics remain stable can an expansion be discussed. Even then, authority or data access is not extended automatically.

The pilot decision

An agent pilot is ready only when a workflow is not expected to fulfil the purpose sufficiently and the mandate, tool permissions, data, memory, approvals, logs, budgets, tests and stop signals are fully described. Otherwise, a deterministic workflow or an assistant is the more controllable solution.

How maitiq helps: from signal to completed task

A helpful agent can read an approved report, detect a deviation and prepare a task with the affected campaign, the supporting data and a proposed action. Whether it creates that task only as a draft or writes directly into a system is a separate decision. For a budget update, the ability to analyse text is not sufficient.

Measure correctly completed tasks, necessary rework and missed exceptions. maitiq supports the design of this workflow and its integration into your team. In the Google Ads product, proposals are traceable and their implementation stays within the recorded authorisation; configured automation can approve and apply within its controls. Human decisions remain important without every implementation being manually approved. Additional agent functions are agreed as part of the project.

How a first pilot with maitiq begins

A scoped discovery phase shows which agent task is really necessary, where a normal workflow is sufficient and which controls a pilot needs. You receive a clearly scoped use case, the open data and control questions, a pilot plan and the criteria for later operation.

Decision rule: Start read-only; afterwards allow only a narrowly limited action with its own authorisation and observable verification.

Sources and how to read them

The four sources have different roles: the OpenAI guide “A practical guide to building agents” describes how agents are built, the role of tools, when a run ends and which controls can be layered; the OpenAI announcement “New tools for building agents” (11 March 2025) supports the definition of agents as it stood then, not current API specifications or prices; the OWASP “Agentic AI Security” initiative adds a narrow security perspective on tool-connected agentic systems; and the FDPIC sets out the data protection framework for processing personal data with AI. Information about how maitiq works is available at maitiq.com.

Have maitiq review your specific case.