•
13 min read

Designing an Agentic AI Flow Inside an Enterprise

Table of Contents

Are dashboards not good enough any more? Why an agent?

We keep a set of fairly plain set-top boxes across a few labs, used for verifying builds and running compatibility tests. When something broke, the old routine was direct enough: an engineer walks over, hooks the box up, runs adb, reads the log, comes back with an answer. And the thing end users report to the platform team most often is exactly this: โ€œsome streaming app will not open on one of the boxes.โ€

This is a situation I run into constantly at work, and I kept wondering whether there was a better way to get to the problem quickly, one where I stop having to worry about: which box is it, is it online right now, which build is on it, where does the data I need actually live. Or worry about how to reach the end user, how to translate their repro steps into the log I need, how to confirm some critical path.

We built dashboards too. But what a dashboard can answer is basically the set of questions somebody already had in mind while building it: is the device online, when was the last heartbeat, where did the log go. It is a panel defined around an assumed scenario, and it is genuinely useful, just not for the situation above, where it cannot get me to root cause quickly. Once you actually start digging, what you need to ask is rarely something you can finish listing beforehand. You might suddenly want to know which schemes an app registered, and there is no button for that on the dashboard, because whoever set it up had no idea this would be needed later.

So what I came to think is that the genuinely useful part of an agent in these situations is not that it can chat, rather that it can compose.

It can list the devices, pick one that is online, pull the log for that window; then, once it sees the log has nothing in it, go again using the callerโ€™s package instead. Each of those steps is a query we use all the time. It just used to take an engineer to string them together before anyone got to an answer.

That need is why I started building an agent to make my own work easier, and to give me my time back.

The first architecture: three layers, and how I came to understand the problem

  somebody asks in a channel
          โ”‚
          โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
   Chat surface                  an LLM agent, and an MCP client
   (Slack bot)                     - re-lists its tools on an interval
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
              โ”‚  carries the asker's identity
              โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
   Gateway                       auth ยท caller attribution
                                 per-tool time budget
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
              โ”‚  swaps in a service identity
              โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
   Domain service                the tools, all read-only (now)
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”˜
       โ”‚               โ”‚
  scheduled        live command
  snapshots            โ”‚
       โ”‚               โ”‚
       โ–ผ               โ–ผ
  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”   โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
    object         Message
    storage        broker
  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜   โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”˜
                       โ”‚
                 โ”Œโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”
                 โ–ผ           โ–ผ
            local agent  local agent
                 โ”‚           โ”‚
                box         box

The most convenient thing, I figured at the start, was to drive this through a Slack bot as the chat surface, with an LLM agent behind it. The bot holds a tool list, and that list is re-fetched on an interval.

So the day we need to add some tool, it shows up on its own without restarting the bot over and over. That looks like a small thing, and once I was actually building it I found it mattered a lot, because getting the behaviour I wanted did mean modifying existing tools and adding new ones fairly often.

The middle layer is the gateway. It is easy at first to think of it as a proxy, and the important job it has ahead of it is being the place identity changes hands. Honestly, I have not made this bot and gateway especially sophisticated yet, nothing like token exchange or a credential broker. What happens today is that the person asking arrives with their own identity, the gateway verifies them, and everything downstream is called with a service identity.

While designing this agent I happened to be reading a lot about agent identity, so I deliberately kept โ€œwho is this service account acting for right nowโ€ around for auditing later. Otherwise every row in the audit log reads as the same robot moving about, and it cannot answer who the work was done for.

At the bottom is the domain service. It has two routes to device data: scheduled snapshots and a live command channel over MQTT. Snapshots are cheap and historical, and they lag. The live channel is precise, and you pay for it with needing the device online and having to handle timeouts.

I did not let the bot call the domain service directly, mostly because tools are only going to multiply, and in the end they will not all be about devices.

Once the gateway layer exists, adding a second domain, whether that is monitoring, CI or releases, is essentially one more source.

Defining the tool description interface

My starting mental model for writing a tool was the one I use for writing function logic: name, parameters, return value, and a description that adds โ€œhere is what this tool doesโ€.

That turned out not to be enough, because the thing actually reading the description is a model. What it has to decide is not only โ€œshould I call this tool nowโ€, but also how far it is allowed to reason once it has that answer.

For example: there is a tool that returns one appโ€™s log on one device, and its description reads like this:

To look into an app that failed to start, you need the callerโ€™s process. Asking with the target appโ€™s package name will not get you the right answer, so ask again with the caller processโ€™s package name, and do not take โ€œno errors under the target appโ€ to mean โ€œno launch failureโ€.

I ended up treating this as an inference rule rather than documentation.

Once it is in there, the model really does carry that boundary out when it answers, and you end up watching the bot tell an engineer in Slack: โ€œan empty handler list does not mean the device answered and nothing claims this URIโ€, and then explain why on its own.

So the ordering I use now for a tool description is simple:

Write what the answer cannot prove first, then write what it can.

The first half matters more, because that is where the model is most likely to fill in the rest by itself.

โ€œThis tool returns heartbeatsโ€ buys almost nothing. โ€œA heartbeat only means the box is phoning home, it says nothing about CPU or memory, because we do not collect thoseโ€ is what actually helps.

How do you verify the modelโ€™s answer?

Putting the rule into the description is only a declaration. Whether the model read it and reasoned from it still comes down to regression tests.

I added an eval suite to this agent. Each case is a JSON file, and it can be used to test and record the quality of the agentโ€™s answers:

incident        the date, what was asked, what the bot got wrong
ground_truth    when it was verified, and against what
prompt          the original question, reproduced as-is
mention[]       regexes that must appear, each with a why
forbid[]        regexes that must not, each with a why
tools_require[] tools the answer must have actually called
judge[]         what a regex cannot decide, handed to an LLM

The part I think matters more is the why field. It does not only explain what the rule guards, it explains where the rule stops. One forbid ruleโ€™s why goes roughly like this:

It deliberately does not match a bare โ€œno errorsโ€. โ€œWe store no errorsโ€ is the honest statement that nothing is recorded, and there is no way to separate it from โ€œno errors were foundโ€, so that judgement is handed to the judge question below.

And another one records that the test itself was changed:

Widened twice. The tool contractโ€™s own word is โ€œpresenceโ€, and a near-verbatim restatement of the correct answer using โ€œhas checked in recentlyโ€ was scored as having missed the evidence.

This is not something I had seen much of in other test frameworks.

A normal test records what the expected value is. This eval tool can keep โ€œthis expected value was wrong before, why it was wrong, and how it was changed afterwardsโ€ alongside it.

I think that matters especially for an agent, because what you are testing is not a deterministic functionโ€™s return value, rather whether a paragraph of natural language said the thing it needed to say. The slightly awkward part is that โ€œwhat counts as the thing it needed to sayโ€ can be written crooked by us too.

Guardrails for letting an agent touch hardware

Once a tool really can reach a machine, guardrails stop being a question of whether, and become a question of how you design and apply them.

Based on the situations I hit most, these are the rules I settled on:

Read-only before agent identity is ready. Anything that changes state requires a human caller; service accounts do not get through for now. Authorization is therefore per action rather than per endpoint.

No generic shell passthrough. The command channel underneath can run any string, reboot included, and the dangerous write operations include deleting data. Until our authority policy is properly in place, only read-only operations are allowed. So each tool assembles its own command, and the caller supplies identifiers and arguments only.

The device-side agent looks roughly like this:

const fullCommand = `${adbPath} ${deviceArg} shell ${command}`;
await execAsync(fullCommand);   // child_process.exec

I did not think much of this at first, until I actually traced what happens to a command after it leaves my side.

exec hands that whole line to /bin/sh -c on the host. Then, when adb shell gets the remaining arguments, it joins them back into a single string with spaces and gives that to the shell on the device to run.

So the same string ends up being expanded twice:

the command string I assembled
   โ”‚
   โ–ผ
the host's /bin/sh -c    # first expansion, eats one layer of quoting
   โ”‚
   โ–ผ argv goes to adb
adb shell                # joins the remaining args back into one string
   โ”‚
   โ–ผ
the shell on the device  # second expansion

I had wrapped the URI in single quotes and considered the matter handled. The problem is that layer of quoting gets eaten on the host, and what the device really receives is the unquoted version.

Later I wrote a stand-in adb script that prints exactly what the device side receives, and the difference is clear:

before โ†’ DEVICE SHELL RECEIVES: ... -d myapp://a$(echo PWNED) ...
after  โ†’ DEVICE SHELL RECEIVES: ... -d 'myapp://aPWNED' ...

The second line is the more interesting one. Wrapping the whole command in another layer of double quotes did keep the single quotes alive all the way to the device, but $(...) had already been expanded on the host, because $ is still substituted inside double quotes. So one more layer of wrapping on its own does not do it.

What I ended up doing is two things together. The wrapping: the whole command goes in double quotes for the host, and the inner single quotes are left for the device. Then the refusal, these five characters are simply not accepted:

CharacterProblemWhy
"the hostโ€™s double quotescloses the outer quoting outright
\the hostโ€™s double quotesescapes out of it
backtickthe hostโ€™s double quotesstill substituted inside double quotes
$the hostโ€™s double quotessame, both $(...) and $VAR expand
'the deviceโ€™s single quotescloses the inner quoting

Everything else a URI carries, ; & | ( ) * ? # and spaces, is inert inside both layers, so it is allowed. Real deeplinks contain these, and refusing them all would leave the tool unusable.

As for feeding it a URI containing $(reboot), the test result is rather good:

Refused, nothing was sent to the device.
uri may not contain a quote, a backslash, a backtick, a dollar sign,
a newline or a control character.

What is not done

today:

  content service โ”€โ”€  sends a link โ”€โ”€โ–ถ device
      โ–ฒ                               โ”‚
      โ”‚                               โ”‚ the tool can read this
   cannot read                        โ–ผ
      โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€  "what format does this app accept"


once the other half exists:

  content service โ”€โ”€โ”
                    โ”œโ”€โ”€โ–ถ reconcile (scheduled) โ”€โ”€โ–ถ mismatch โ”€โ”€โ–ถ Slack
  device snapshot โ”€โ”€โ”˜

The content side. These tools can answer what an app on a device accepts. They cannot answer what we sent it. That one is structurally unavailable on the device, because the platform writes the URI into the log through a safe-string pass that keeps the scheme and host and replaces the path with /....

Reconciliation. Once both halves exist, drift can be found before a user reports it. Right now we are still reactive, going to look only after somebody tells us.

Live probe versus scheduled snapshot. Every tool that touches a device today is live, so it needs the device online. The other route is to have the local agent upload an inventory on a schedule and let the tool read that. They complement each other and I want both: snapshots for everyday questions, live probes as ground truth.

Right now, if I want to know which agent version is on a particular box and which commands it supports, I still have to ask a person. That alone says the system is missing a piece of observable state.

It should report itself, the same way a device reports its build.

This system has not been in production long, so how well it works and how much time it actually saves is still to be seen. There are probably issues I have not found yet: which descriptions are still too thin, which tools the model keeps picking wrongly, which tools nobody calls at all. There is simply not enough usage yet.

That part only comes once real traffic does.