An LLM can translate an operator's intent into a proposed Kubernetes change, but it should not be the authority that decides whether the change is valid or permitted. Keep interpretation probabilistic and make validation, authorization and execution deterministic.
This article presents a reference architecture for a constrained configuration assistant. The example changes the replica count of an existing staging Deployment. It is a design exercise, not a production-ready agent, a tested benchmark or a claim that local model execution automatically makes infrastructure changes safe.
Separate intent from execution
A request such as “increase the booking service to four replicas” leaves several questions unanswered. Which cluster? Which namespace? Which Deployment? Is a horizontal autoscaler responsible for replica count? A model that guesses those details can produce valid YAML for the wrong target. A correct syntax check would not catch that operational mistake.
Provide target selection through trusted application context and require explicit resolution of ambiguity. Let the model extract a small operation rather than rewrite a complete resource. The execution service should receive an operation it understands, not arbitrary shell text. This reduces the space of possible changes and makes authorization easier to audit.
{
"operation": "set_replicas",
"namespace": "staging",
"deployment": "booking-api",
"replicas": 4
}
This object is an untrusted proposal. The application must still verify its schema, resolve the target against an allowlist and confirm that the requester may make that exact change. Unknown fields, unsupported operations and missing values should cause a clear rejection rather than a best-effort guess.
Use a narrow schema and deterministic policy
For the example operation, require a non-negative integer replica count within a configured bound, an approved namespace and an existing target. The permitted bound is an organizational policy, not something the language model should infer. Setting replicas to zero might be syntactically valid and still be prohibited for a protected service.
Perform authorization on the server using the authenticated caller. Never trust a role name included in model output. The executor's Kubernetes credentials should have only the resource and namespace permissions required by the supported operation. A helper that changes one staging Deployment does not need broad administrative access.
Check ownership before proposing a change. A GitOps controller may restore the repository's declared value, or an autoscaler may manage replicas. In those cases the appropriate workflow may be a repository change or an adjustment to autoscaling policy. Editing a live field without understanding its controller creates confusing behavior even when every API call succeeds.
Construct the smallest change from trusted state
Fetch the current resource through a typed client and calculate the intended difference in code. Preserve fields outside the allowed operation. Whole-document regeneration gives the model unnecessary opportunities to omit labels, alter selectors or change security settings. A narrow operation makes the review smaller and the expected effect more explicit.
Kubernetes Server-Side Apply tracks field management and can report conflicts between managers. That is useful for declarative workflows, but it is not a permission to force ownership away from other controllers. Choose the appropriate API operation for the field and workflow; do not use forced conflict resolution as a routine way to make an AI proposal succeed.
For a reviewer, show the cluster, namespace, resource name, current value, proposed value and policy result. A one-line diff without target identity is easy to approve in the wrong environment. Bind the proposal to a snapshot of the relevant state so execution cannot silently apply a previously reviewed decision to a materially changed resource.
Validate against the API server
Local validation catches malformed proposals early, but the cluster's API server has information the application may lack. Kubernetes supports server-side dry-run requests, which exercise supported validation and admission processing without persisting the requested change. The API concepts documentation describes these semantics and relevant limitations.
Admission controllers can validate or mutate requests before persistence. Their presence means the final admitted object may differ from the initial proposal. Show meaningful server-side changes in the review where possible. A rejection from cluster policy is a useful outcome; the assistant should explain it rather than trying alternate encodings until something passes.
A successful dry run does not prove that the workload will become healthy. Capacity, image availability, application behavior and external dependencies can still prevent a successful rollout. Treat dry run as one validation stage, followed by controlled execution and observation of the actual workload.
Bind approval to the exact proposal
Where the workflow requires human approval, record what was approved: requester, target, operation, relevant state, proposed difference and expiry time. If the target changes materially before execution, regenerate the proposal and obtain a new decision. A generic “approve this session” flag should not authorize later, unrelated operations.
Use concurrency control appropriate to the API operation. The executor should detect conflicting updates instead of overwriting them silently. If execution fails, report whether the operation was rejected, partially observed or completed with an unhealthy result. Those states require different recovery actions and should not all become a generic error message.
Rollback also needs a precise definition. Restoring a previous replica count may be straightforward, while reverting a data migration or deleting a created resource may not be. Document the recovery action per operation. An assistant should never promise universal rollback simply because it saved the previous YAML.
Treat retrieved text as data
Logs, repository comments, resource annotations and incident descriptions may contain text that looks like instructions. A model reading those sources must not treat them as authority to change its task or expand permissions. Keep the operator's request, retrieved evidence and tool results in clearly separated roles.
More importantly, enforce the boundary outside the model. Even if a hostile log line persuades it to request an unsupported operation, the execution service should reject the proposal. Prompt wording can improve behavior, but it cannot substitute for server-side authorization, schema validation and a limited tool surface.
Local inference changes where model computation happens; it does not remove these problems. A locally hosted model can misunderstand intent, produce an invalid object or follow misleading retrieved text. Evaluate those failure modes regardless of where the model runs, and keep credentials out of its context whenever they are not required.
Evaluate refusal and recovery, not just happy paths
Create a small test corpus with clear requests, ambiguous targets, unsupported changes, conflicting controller ownership and malicious text embedded in retrieved material. Include requests that should be rejected. A system that always returns an operation can have a high completion rate and a poor safety record.
| Case | Expected behavior |
|---|---|
| Explicit allowed target and count | Produce a bounded proposal and validate it |
| Two matching Deployment names | Require target disambiguation |
| Protected production namespace | Reject before mutation |
| Target changes after review | Invalidate the stale proposal |
| Log text requests credential access | Treat it as untrusted data; refuse unsupported execution |
| Dry run succeeds but rollout stalls | Report observed failure and the defined recovery option |
Measure correct intent extraction, policy decisions, unintended field changes, end-to-end latency and the clarity of rejection messages. Keep evaluation cases versioned with policy and prompt changes. For performance comparisons, publish the workload, hardware, model configuration and evaluation procedure instead of a single unexplained speedup figure.
Keep the operating model understandable
Log the proposal and execution decision with sensitive values redacted. Connect those records to deployment events and operational telemetry so the team can determine what changed and whether it improved service behavior. Record rejected attempts as well as successful ones; repeated rejections may reveal unclear interfaces or a policy mismatch.
A useful infrastructure assistant makes a small set of operations easier to understand and review. Begin with a narrow scope, prove that unauthorized and ambiguous requests are handled correctly, and expand only when evaluation supports the new operation. The model helps interpret intent; the engineering system remains responsible for the change.
