By John Wheeler
Companies are getting very good at giving AI tools a task. They are less sure what to do when the tool starts doing that task every day.
In a demo, an AI agent can draft a report, answer a customer question, or sort through a pile of records in seconds. That can be impressive. But once the agent becomes part of daily work, a different set of questions comes up. Who owns the result? What information may the agent use? What is it allowed to change? How do we know its work is good? What happens when it gets something wrong?
Those questions sound a lot like management questions. That is why I think an AI agent needs something familiar: a job description.
I do not mean a long document written to satisfy a policy. I mean a short, practical agreement about the work the agent is expected to do and the limits around it. If a business cannot describe the job clearly, it will struggle to judge whether the agent is doing it well.
Start With the Result
A human job description usually begins with a role. An AI role should begin with a useful result. “Help with marketing” is too broad. “Draft three short posts from approved research and place them in a queue for review” is clearer. We can tell whether the work happened, whether the drafts were usable, and how much review they required.
The result should connect to a business need. A faster draft has little value if someone must spend more time fixing it than they would have spent writing it. A detailed report has little value if it misses the one fact needed to make a decision. When I evaluate machine work, I want to know what was accepted, what needed correction, and what changed because of the work.
This also keeps us from confusing activity with progress. An agent can run all day, produce pages of text, and still add little value. A smaller job that saves a manager one hour of careful work may be the better use of AI.
Name a Human Owner
“The AI did it” is not an acceptable management answer. Every AI role needs a human owner who can explain its purpose, judge its results, and decide when the role should change.
That owner does not need to read every line forever. In fact, a good system should reduce needless review as it proves itself. But someone must be responsible for the outcome. If the agent publishes a false statement, changes a key record, or overlooks an exception, the company needs to know who can investigate and fix the process.
Ownership also gives the agent a place in the organization. A worker preparing a financial report should know whose rules govern the report and who receives an exception. A worker drafting product content should know who approves a claim before it goes public. Without that link to human judgment, a collection of agents can become a collection of loose ends.
Set the Information Boundary
The next part of the job description is the information the agent may use. This can include approved documents, selected records, public sources, or a limited set of tools. It should also say what the agent must leave alone.
Access is easy to overlook because many AI demos begin with a person copying information into a chat window. In daily operations, the agent may read from several systems and combine what it finds. A company should be able to answer a plain question: Why does this role need access to this information?
Good access rules help with quality as well as risk. If an agent uses an old price list or a draft policy, it may give a polished answer built on the wrong facts. The role needs a source of current, approved information and a way to flag conflicts. “I found two different answers” can be a useful result. Quietly choosing one is often a problem.
Make Authority Specific
The word “autonomy” can hide several very different actions. An agent might be trusted to read a document, draft a recommendation, update an internal record, send a message, or make a purchase. Each action has a different effect on the business.
That is why authority should be granted by action. An agent may draft a customer reply while a person decides whether to send it. It may spot an inventory problem while a manager decides how to resolve it. After repeated good performance, it may be allowed to handle a narrow, routine case on its own.
The boundary should include a stopping rule. If the agent finds missing facts, conflicting instructions, or a case outside its normal scope, it should ask or hand the work to someone with authority. Knowing when to stop is part of doing the job well.
Test the Ordinary Tuesday
AI demos often show the best case. Real work includes routine cases, awkward cases, and cases no one expected. The true test is an ordinary Tuesday: Can the agent do the usual work, notice when something is unusual, and leave a clear record of what happened?
I would start with a narrow task and a set of examples that reflect actual work. Some should be straightforward. Some should contain gaps or mixed signals. Then I would compare the agent’s work with a clear standard. Did it produce the right result? Did it cite the right source? Did it ask for help at the right time? Did it take an action it was not allowed to take?
A single good answer is encouraging, but it is not proof of reliability. Repeat the test. Change the wording. Use fresh examples. Watch what happens when the information changes. A role earns more trust by showing steady results over time.
Measure the Cost of Review
Many teams count how much work an agent produces. I would also count the work it sends back to people. How long does a reviewer spend checking it? How often does a manager have to step in? Are exceptions clear enough to decide quickly, or does the person have to reconstruct the agent’s entire path?
Human attention is a real cost. If every agent routes every decision to the CEO, the company has made a faster path to CEO overload. The answer is to define who owns each type of decision, give reviewers the evidence they need, and keep routine matters at the right level.
This is also why an error rate alone is not enough. A minor drafting error caught in review is different from an incorrect change to a customer record. The same agent may be reliable enough for one job and not yet reliable enough for another. Measure the effect of a mistake, not just the number of mistakes.
Turn Failures Into Better Work
When an agent fails, the first question is what happened. The second is what should change. Was the instruction unclear? Was the source out of date? Did the role have too much authority? Did the test miss a common case? Did the human reviewer lack the right evidence?
A useful lesson should lead to a visible change: a clearer rule, a better source, a new test, a revised approval step, or a narrower scope. Then the next run can show whether the change helped. Saving a note about the failure without changing behavior is not much of a learning loop.
This approach also gives people a fair way to judge the technology. An agent does not need to be perfect to be useful. It needs a job where its strengths create value, its mistakes can be found, and the business can respond before those mistakes grow.
Let Trust Grow in Stages
I would not give an agent a broad title and hope it acts responsibly. I would begin with a defined job, limited access, and a human reviewing the results. As the role proves itself, I would expand its authority one action at a time. If results weaken, I would narrow the scope, fix the process, and test again.
That is how trust works in a well-run company. We do not give a new employee every permission on the first day. We explain the work, provide the right tools, review the results, and grant more responsibility when it is earned. AI roles deserve the same care, even though they work at a very different speed.
The practical question for leaders is not “How many agents can we launch?” It is “Which job can we define well enough to measure, improve, and trust?” Start there. A clear job description turns an impressive demo into work a company can actually manage.
