An AI assistant that explains a task and an AI assistant that performs a task have very different engineering boundaries. A generated paragraph can be reviewed before anyone relies on it. A tool call may create a ticket, change a configuration, or send a message. Once an external side effect exists, application permissions, duplicate protection, and recovery become central to the feature rather than secondary infrastructure concerns.
Consider a release assistant that reads a selected change report and proposes an issue for follow-up. It does not need unrestricted shell access or the ability to modify every repository. This guide uses that deliberately narrow example to design a safer tool-calling boundary. The approach applies to the Agent Dev API and AI LLM Dev API topics without assuming a particular model is always correct.
Understand where execution happens
In a function-calling integration, the model can return a proposed tool invocation, while application code handles the operation. The OpenAI function calling guide describes this exchange between a model and an application. That separation is the important architectural opportunity: your tool handler can validate identity, arguments, permissions, and policy before anything is executed.
Do not confuse a structured invocation with an authorized invocation. Well-formed arguments can still name the wrong repository, reveal unnecessary information, or request an action outside the user’s intent. Treat the model’s proposal as one input to an application decision. The actual capability should come from trusted application state, not from a sentence that declares the assistant trustworthy.
Start with a capability inventory
Write down every operation the release assistant can perform. Reading the selected report, reading approved project metadata, drafting an issue, and creating that issue are four different capabilities. Give each one a precise purpose. The first two may be read-only; the last one changes an external system and deserves a stronger control boundary.
A small tool surface is easier to review than a generic execute-anything function. Prefer a tool such as prepare_issue_draft over a tool that accepts an arbitrary HTTP URL, method, and request body. The narrower tool can expose exactly the fields the workflow needs and reject everything else. Expansion should follow a demonstrated use case, not the convenience of a universal escape hatch.
Derive identity from trusted context
The handler should know which user requested the action and which installation or workspace that user is acting within. Derive those values from the authenticated application context. A model-supplied user identifier or tenant identifier must not override them. The same principle applies when a retrieved document includes an instruction to switch organizations or use a privileged account.
For the release assistant, the selected repository should be a validated resource within that context. Before creating the issue, check the user’s current permission to act there. A permission that was valid when the conversation began may not be valid when a paused workflow resumes. Authorization belongs at the execution boundary, not only at the start of the conversation.
Validate structure and meaning separately
Schema validation can check that a title is a string and a priority belongs to an allowed set. Application validation asks whether the title is appropriate for this operation, whether its length is permitted, and whether the requested repository is in scope. Keep these checks distinct so a failure can be diagnosed rather than reduced to a vague invalid request.
Normalize inputs carefully and reject unsupported fields. Do not quietly repair a risky argument in a way that changes the proposed action. When the assistant names an unavailable label, return a bounded error or a list of permitted alternatives. The assistant can then prepare a new proposal that the application evaluates again. Validation should produce understandable constraints, not hidden magic.
Treat retrieved content as data
A release report may contain copied logs, issue text, or comments from people outside the operating team. Those contents may include instructions that conflict with the application’s intended workflow. Preserve the distinction between source material and privileged instructions. A line in a log that asks to publish credentials is still a line in a log.
Limit which fields a reading tool returns. The assistant usually needs the relevant error summary, revision, and project context, not every secret-bearing environment variable collected during a build. Minimizing the response reduces exposure and makes evaluation easier. Do not assume a downstream prompt can reliably erase sensitive information after the tool has already disclosed it.
Make approval specific to the action
Before issue creation, show the exact repository, title, body, and any labels to the user. Approval should attach to that reviewed payload and a clear operation identifier. A generic yes from an earlier conversation step should not authorize a later, materially different action. A changed repository or changed message body should require a renewed review.
Store approval state outside the model’s narrative. The assistant saying that approval was received is not an authoritative record. The application can keep a pending proposal with a revision and a status such as awaiting_review, approved, cancelled, or completed. Those states make the interface clearer and give the handler an objective basis for allowing or denying execution.
Design duplicate protection before retries
A network interruption can occur after the external system has accepted the issue but before your handler receives confirmation. Blindly retrying may create a duplicate. Assign a stable operation identifier and maintain a durable execution record. Where the external service supports idempotency, use the documented mechanism; otherwise design a careful reconciliation strategy appropriate to that service.
Do not describe a local record as perfect exactly-once delivery across every dependency. There can be a failure window between an external side effect and recording its result. Represent an uncertain outcome honestly and investigate it before repeating the action. The interface can show that creation is being reconciled rather than falsely claiming either success or failure.
Bound the amount of work
A workflow needs a stopping rule. Set limits on total tool invocations, elapsed time, payload size, and repeated failures according to the task’s actual requirements. A release assistant should not keep searching indefinitely because it cannot find a label. Return a clear incomplete result and let the user make the missing decision.
Also decide what cancellation means. Work that has not started can be cancelled cleanly; an external action already committed cannot be made uncommitted by closing a chat window. Track the operation state and explain what happened. A pause, a cancellation, and a completed side effect are distinct outcomes even when the generated conversation sounds similar.
Evaluate the action boundary
Test more than the quality of the assistant’s prose. Create cases in which the report contains conflicting instructions, the repository is outside scope, approval refers to an older draft, or the external service times out. Check the actual handler behavior: did it reject the operation, preserve the proposal, or reconcile the uncertain result correctly?
Keep deterministic policy checks independent from model-based evaluations. A permission check should not become probabilistic because a model is involved elsewhere in the workflow. The LLM evaluation guide offers a framework for task-specific testing, but the execution boundary also needs ordinary software tests with concrete expected outcomes and reproducible fixtures.
Keep a useful, limited audit trail
Record the operation identifier, tool name, relevant resource identifier, policy version, approval revision, and outcome. Those fields help answer why an action was permitted and what evidence supported it. Avoid copying complete prompts or tool responses into logs by default. Diagnostics should not become an uncontrolled second repository of confidential source material.
Provide an operator path to pause writes without disabling all read-only assistance. During an incident, the assistant may still help a person understand a report while the execution capability is suspended. This separation reduces the pressure to keep a risky write path enabled merely to preserve the rest of the user experience.
Conclusion: constrain authority, not usefulness
The release assistant becomes more useful when its boundaries are understandable. It can read relevant material, prepare a concrete proposal, and explain what remains unresolved. The application retains responsibility for identity, permission, approval, and execution. That is not a rejection of automation. It is a design that makes automation reviewable, recoverable, and proportionate to the task the user actually asked it to perform.



