LLMs Need Context Before They Can Help IT Operations
Series: From Automation to Autonomous IT Blog 5 of 11
Quick summary: Useful LLM assistance starts with a bounded, permission-aware view of the current environment, not a more elaborate prompt.
TL;DR: Give the model an evidence packet with scope, timestamps, relationships, operating constraints, and explicit gaps. Ask it to explain what that evidence supports and what to check next. Keep operational truth in the source systems and production authority in governed orchestration.
An operator asks a large language model (LLM) why a payment service is failing. The answer points to connection-pool exhaustion and recommends restarting the application tier. It even cites a familiar runbook.
It sounds reasonable. But does the model know which payment service is affected? Has it seen the current dependency map? Is the runbook for this deployment? What changed before the failures began?
Until those questions have answers, the response is a plausible explanation of a class of problems, not a diagnosis of this incident.
In the previous article, we asked what an operator needs to act on a group of alerts. Service impact and a clear response owner matter more than a shorter alert list. That actionable event is a useful starting point for language assistance. The model can help people interpret the record. It should not fill missing evidence with a confident story.
The leadership question is therefore narrower than “Which model should we buy?” It is: “What can this assistant actually know about our environment when someone asks for help?”

A runbook is not a view of the current environment
A general model may explain common failure patterns. An internal document collection may describe the organization's procedures. Neither should be treated as evidence of the service's current state unless a current, relevant observation is actually supplied.
Retrieval-augmented generation (RAG) combines model generation with retrieved information. The foundational RAG paper describes a model that uses both learned parameters and an external retrieval index.[1] That architecture is useful background, not proof that a particular implementation has fresh enterprise data or the right access controls.
A runbook can explain how to investigate a failed connection. It cannot establish that this connection is failing now. A historical incident can offer a diagnostic lead without showing that today's incident has the same cause.
For operations, design the context path to combine relevant documents with authorized observations from the systems that own the facts. The assistant needs approved procedures alongside current observations. Relevant changes and dependencies help establish whether a procedure applies. Label them differently so the assistant can distinguish an instruction from an observation.
Do not make the model the place where infrastructure state becomes authoritative. An operator should be able to follow a reference back to the source record and see what was known at the time of the answer.
Build a context contract for one operational question
Start with a task you can evaluate, such as preparing an incident briefing for one service. “Understand our entire infrastructure” is too broad to be a useful first acceptance test.
For that task, define a context contract. Specify what the assistant needs to know and is permitted to see. Decide how it should respond when required evidence is unavailable. This is a practical design checklist, not a standard schema or a claim about native product fields.
Five questions keep it concrete.
1. Scope: what are we talking about?
Identify the service, environment, and relevant assets using stable identifiers. Include the requesting user's role and the purpose of the request.
A service name without an environment can mix production with a test deployment. An asset nickname can point at a retired system. Resolve those ambiguities before retrieval, or have the assistant ask for clarification.
Apply access restrictions before supplying the model with data. The answer must not reveal another team's restricted incident, credentials, or customer information simply because it matches the query. “Read-only” describes execution authority; it does not make unrestricted data access harmless.
2. State and time: what was observed, and when?
Give each relevant observation its observation time. A retrieval timestamp alone cannot tell you how old the underlying evidence is. Identify the time window being compared.
Freshness should match the task. A configuration baseline may remain useful longer than a connection count during an active outage. Define that tolerance for the workflow rather than applying one universal age limit.
Do not let the assistant describe observations from different times as a single current snapshot. If a collector has stopped reporting, record the gap. A retrieval failure is not evidence that the service is healthy.
3. Relationships and change: what could connect the symptoms?
Include the dependencies relevant to the affected service and the changes that could plausibly affect them. Preserve the source and age of the relationship data.
This builds on the infrastructure digital twin discussed earlier in the series: an operational representation of assets, relationships, and state. It need not be a complete replica of the enterprise to help with a bounded question.
A shared dependency or a recent change can narrow an investigation. Neither proves causation. Ask the assistant to name the evidence that supports a hypothesis and the observation that would weaken it.
4. Operating constraints: what response is available?
Supply the applicable runbook version, accountable owner, and relevant restrictions. A maintenance window, change freeze, approval requirement, or unavailable recovery path can change the appropriate next step.
These constraints help the model avoid proposing an obviously unsuitable response. They do not authorize it to carry out a change. The execution process must independently enforce permissions and policy against the state that exists when action is requested.
If the question is about capacity or placement, cost and locality may matter. For a diagnostic summary, they may not. Include context because the task needs it, not because the data exists.
5. Evidence gaps: what remains uncertain?
Expose evidence gaps. Flag stale records and disagreements between sources. Separate direct observations from inferred relationships and working hypotheses.
NIST's Generative AI Profile identifies confidently stated false or erroneous content as a risk it calls confabulation.[2] In an incident briefing, a useful response should expose uncertainty rather than smooth it into a narrative.
“The topology record is older than the latest change” is more helpful than an unexplained confidence percentage. Define when the assistant should narrow its answer, request a read-only check, or return the task to the operator.
A hypothetical incident: the same alerts, a different briefing
Consider a hypothetical payment-service incident. This example illustrates an operating model, not a deployed customer result or measured improvement.
Application monitoring reports failed transactions. Database clients report timeouts. A network tool reports a link failure. An operator asks whether the application tier needs a restart.
With only those alert summaries, restarting is one possible response among several. The assistant has no basis for treating it as the right one.
Now give it a bounded evidence packet: the affected production service and observation times, a dependency record showing a shared upstream path, and a change record for that path preceding the failures. Add the approved read-only diagnostic procedure and note that the relationship map has not been refreshed since the change.
Ask for a briefing, not a fix:
Summarize the observed impact. Separate facts from hypotheses. Cite the supplied records, identify contrary or missing evidence, and propose the next permitted diagnostic check. Do not infer that a production change is authorized.
A useful briefing would identify the shared path as a candidate fault domain while making the stale map explicit. It would explain why connectivity should be checked before treating the application as the cause. It should not claim that the network caused every timeout or that restarting would restore service.
If a fresh check shows the payment service no longer uses that path, the hypothesis should change. The evidence must be able to correct the answer.
That is where language assistance belongs: helping a responder understand an incomplete record and choose a defensible next investigation step. Urgent service response must continue even if the model or context service is unavailable.
Evaluate the briefing before expanding its authority
Choose a narrow initial role, such as preparing incident briefings for one service. Runbook retrieval or diagnostic proposals could follow once that role is reliable. Keep consequential production actions out of the pilot while you assess whether the assistance is useful.
Use reviewed incidents and controlled exercises. Test stale topology and simultaneous unrelated failures. Then remove a data feed or supply conflicting records. Include a runbook for the wrong environment. Check restricted-data handling separately; a fluent answer is not evidence that permissions were respected.
Evaluate each response against the evidence available at that moment, not against facts discovered later. An incident's eventual root cause should not leak into a replay of its initial briefing.
For one workflow, track:
Whether material factual statements are supported by the referenced records.
Whether operators can retrieve those records and inspect the relevant evidence.
How often required context is unavailable or outside its freshness tolerance.
Which proposals operators accept, correct, or reject, and why.
The time needed to assemble and review a usable briefing, including corrections.
Keep the pilot's evaluation records permission-scoped too. Do not create an unrestricted archive of sensitive prompts and retrieved material in the name of auditing.
If incident performance improves, preserve the measurement boundary. Mean Time to Identify (MTTI) runs from confirmed detection to identification of an actionable cause or fault domain and the affected service. Finishing a summary does not end that interval. Mean Time to Recovery (MTTR) runs from the beginning of service impact to confirmed restoration of the service's agreed operating condition. An accepted recommendation is not a recovery result.
Improve the evidence path before adding more autonomy
When a briefing is wrong, inspect where it went wrong. Did retrieval select an obsolete runbook? Did identity resolution choose the wrong environment? Was the source contradictory, or did the model make a claim the source never supported?
Those failures require different fixes. A stronger model does not repair a missing feed. More documents do not resolve ambiguous asset identities. A more persuasive answer does not establish that a proposed action is permitted.
For leaders, the first deliverable should be one workflow’s context contract, backed by tests that expose the assistant’s limits. Expand the scope only after the team can inspect the evidence, recognize gaps, and correct the response.
Context makes LLM assistance more useful. It does not make production action authorized. The next article, The Enterprise Guardrails That Make LLM-Assisted Operations Trustworthy, examines the controls that must remain in place when a proposal moves toward execution.
If your organization is working through these questions, Orchestral.ai can help you assess one cross-domain workflow and identify the visibility, policy, orchestration, and recovery capabilities needed to move forward. That focused assessment can establish a practical baseline for the next stage of your operating model.