AI & DETECTION · 11 MIN

Sentry Workbench : A Human Led AI Investigator .

Sentry is the investigation console we designed for our own SOC analysts. AI drafts a cited verdict. A human decides, edits and signs every one.

Gregori Nazarovsky· Cofounder & CTO, QMasters· 2026-10-11
TL;DR

What is a human led AI SOC, and how does AI investigate without replacing the analyst?

A human led AI SOC uses AI agents to gather evidence and draft investigations, while analysts decide what gets investigated, check the work and own the verdict. Sentry, the investigation console QMasters designed for its own SOC analysts, enforces that split in its architecture. Nothing is investigated until an analyst requests it. Every line of the AI draft cites evidence and is labeled as fact, context or inference. Hypotheses are logged before their tests run. Every customer boundary is enforced below the AI. No draft becomes a verdict until a named analyst signs it. Sentry is designed and prototyped, and the build is starting.

The Sentry workbench showing an AI drafted investigation under human review, with every timeline row citing its evidence
The Sentry workbench showing an AI drafted investigation under human review, with every timeline row citing its evidence

Sentry Workbench : A Human Led AI Investigator .

By Gregori Nazarovsky, Cofounder & CTO, QMasters

One line sits at the top of every investigation in Sentry.

The human decides and owns the verdict.

Everything else in the architecture exists to make that line true.

Sentry is the investigation console we designed for our own SOC analysts.

It is where our Human-Led AI SOC meets a real case: an AI agent gathers the evidence and drafts the investigation, and an analyst decides what it means.

This post walks through the architecture, and the reason behind each choice.

*Sentry is designed and prototyped, and the build is starting now.

The screens in this post are the design prototype running on invented demo data.

Every company, person, address and ticket in them is fictional.*

The short version

What it is.

An internal console for QMasters SOC analysts.

Every ticket from the customer portal appears in Sentry under its customer.

When an analyst asks, an AI agent investigates that one customer's case and drafts a cited verdict.

The analyst steers the work, accepts, edits or rejects each section, and signs.

What it is not.

It is not an autopilot.

The machine never picks its own work, never runs a query on its own say so, and never turns a draft into a verdict.

What the architecture enforces.

Every factual line cites evidence from the case.

Every line is labeled as a fact, context or an inference.

Hypotheses are logged before their tests run.

Every customer boundary is enforced below the AI.

Every write action waits for a person.

Where it stands.

The design is complete and the screens are drawn.

The first milestone is not a launch.

It is a blind grade of the AI's drafts by our senior analysts.

The Sentry workbench: the analyst queue on the left, the AI draft under review in the middle with every timeline row citing its evidence, and an Ask answer on the right next to a proposed query that has not run

*Figure 1.

The Sentry workbench: the analyst's queue, the AI draft under review with every timeline row citing evidence, and an Ask answer built only from evidence already in the case.

The proposed query on the right has not run, and will not run until an analyst runs it.

Design prototype, invented demo data.*

Where Sentry sits in a Human-Led AI SOC

Our model has three layers.

In the execution layer, agentic AI works at machine speed.

In the judgment layer, senior analysts validate the signals that carry real impact.

In the outcome layer, the work lines up with each customer's own business risk.

Sentry is the judgment layer's workbench.

It is where machine speed meets human judgment, one case at a time.

At QMasters, this design supports our human-led SOC and StrongHold MCSS. For the wider operating model, explore Agentic SOC.

We did not want a faster way to produce confident text.

We wanted a faster, better evidenced first draft that a person owns.

That difference drove every decision below.

1. The machine never picks its own work

Every ticket from a customer's portal shows up in Sentry, under that customer.

None of them is investigated until an analyst presses Process.

There is no auto triage and no background investigation anywhere in the design.

Each Process spends one credit from the analyst's monthly allocation.

One credit buys one full investigation of one ticket.

The point is not to ration the tool.

It is to see, analyst by analyst, how much the team leans on it.

A queue that investigates itself looks efficient right up until nobody can say who decided what.

In Sentry the answer is always a name.

2. Every line is a fact, context or an inference

The AI's draft reads like an analyst's report, with one difference.

Every line carries a label.

A fact cites the evidence row it came from, such as EV-0004.

Context is useful but is not evidence, such as how the same entity and the same rule closed over the past 13 months.

An inference is the AI's reasoning, marked as reasoning.

The labels stop the most common failure in AI writing: a guess delivered in the same voice as a measurement.

History can route an investigation and seed a hypothesis.

It never carries a verdict on its own, and the screen says so in a banner: context, not evidence.

The verdict carries its own honesty.

In the example below it reads needs more, ask the customer, confidence low, and the word uncalibrated sits next to the confidence on purpose.

Until we have graded enough cases to calibrate it, a confidence number is an opinion, and the screen admits it.

Every section of the draft has three buttons: accept, edit, reject.

Nothing is decided until an analyst decides it.

The draft stays unsigned until a named owner signs it.

Behind the labels, a citation gate checks the work.

The cited evidence must exist in this case, and any number in the claim must appear in that evidence.

A claim that cannot point at evidence does not reach the analyst as a fact.

The draft report in dark mode with every summary line labeled fact, context or inference, a verdict of needs more with confidence marked uncalibrated, and the hypothesis ledger beside it

*Figure 2.

Every line of the draft is labeled fact, context or inference, and the confidence is marked uncalibrated.

The hypothesis ledger shows each test registered before its result, a refuted hypothesis kept on the page, and a test that did not run.

Design prototype, invented demo data.*

3. Hypotheses are logged before the evidence arrives

An AI that sees the answer first can always write a test the answer passes.

So Sentry keeps a hypothesis ledger.

Each hypothesis, such as compromised credentials or a travelling user, is registered with a timestamp.

Each test that could support or refute it is registered with a timestamp too.

When a result comes back, the ledger checks that the observation post dates the registration.

A test written after the fact is visible as one.

Refuted hypotheses stay on the page.

In the example, an apparent device clock problem was tested and refuted, and it stays in the ledger marked refuted, because what we ruled out is part of the investigation.

Tests that did not run stay visible too.

In the example, the AI reached its step limit before one query ran.

The ledger marks that test not run and offers it as the cheapest check that could flip the verdict, waiting for an analyst to approve it.

Gaps are never silent.

A failed query is recorded as a gap in coverage, not dropped.

A case whose data source is degraded parks and says so, instead of carrying on half blind.

And thin coverage has exactly one outcome.

If the evidence a playbook expects is missing, the verdict is forced to needs more.

It is never allowed to drift to benign.

4. The graph says how it knows

Investigations are about relationships.

This account signed in from that address, that address got a VPN lease, and the lease reached this server.

Sentry draws them as a graph, and every edge carries how it is known.

A solid line is observed activity, and clicking it opens the evidence.

A dashed green line is approved on record.

A dashed grey line is probable, not confirmed.

A dotted amber line is inferred, and an inferred link is a proposal, not a fact.

In the example, the AI suspects that a second account, an admin variant of the same user name, belongs to the same person.

The graph shows that link as a candidate with its confidence.

It never merges the two identities on its own, because a low confidence identity match must never surface as fact.

The graph can show the picture as of the alert or as of now.

When it expands to neighbors of neighbors, it follows only rows the case already cites.

It never expands a supernode, such as a group that every user belongs to.

The entity graph around one user, with edges drawn solid for observed activity, dashed green for approved on record, dashed grey for probable, and dotted amber for an inferred link to a candidate second account

*Figure 3.

The entity graph draws how each connection is known: observed, approved on record, probable or inferred.

The suspected second account is a candidate, not a merge.

Design prototype, invented demo data.*

5. Ask answers from evidence, and new queries wait for a person

Analysts can ask the case questions in plain language, such as what we know about an address.

The answer is built from evidence already collected, with each piece cited, and it says when no new query ran.

When an answer needs more data, Sentry proposes the query and shows its exact text.

It does not run it.

A proposed query never runs on click, it spends from a visible run budget, and it needs an explicit run step.

The budget exists for a practical reason.

Reads are not free.

A burst of heavy searches can slow a shared SIEM for every customer on it, so the design caps concurrency and trips a circuit breaker before curiosity turns into an outage.

Write actions sit behind a stricter rule.

They are gated by their side effect, not by their name.

A script that calls itself read only but runs on a customer's endpoint counts as a write, and it waits for a person with the authority to approve it.

Every call, every approval and every refusal lands in a hash chained audit log, anchored off the server so that a deleted record cannot hide.

6. One case, one customer, enforced below the AI

Our analysts work across more than 240 customer environments.

In that setting the most dangerous AI failure is not a wrong verdict.

It is the right answer about the wrong customer.

So in Sentry every case belongs to exactly one customer, and keeping that boundary is not the AI's job.

The AI never names a customer in a tool call.

The binding between a case and its customer is held on the server.

A scoping broker, the only path from Sentry to any customer system, resolves that binding on every call.

The broker rechecks the analyst's permissions on every call as well, so access revoked in the middle of an investigation stops in the middle of the investigation.

Underneath, each source is reached with access scoped to that customer at the platform itself.

If the scoped access is missing, the call fails closed instead of falling back to anything broader.

Binding a customer to its sources takes two people.

A daily attestation probe then checks that each customer's access is as narrow as it is supposed to be, because a scope that was wrong on day one is the breach no test suite will catch.

The screen states the rule in one line: Sentry never mixes tenants, and switching cases relocks the data boundary.

The first milestone is a grade, not a launch

Architecture can make an AI's work checkable.

It cannot make it good.

So the first step of the build is deliberately disposable.

Before we build the platform around it, the AI investigates real past cases, our senior analysts grade the drafts blind, and the result is a hard go or no go.

If the drafts are not useful, no amount of isolation and audit machinery will save them, and we would rather learn that cheaply.

After that, every change to prompts, playbooks or models is replayed against a frozen set of past investigations before it reaches an analyst.

We will also measure the people, not just the AI.

A draft accepted without a single edit is the result we audit most, because rubber stamping looks exactly like success.

A verdict an analyst flips after digging deeper is the most valuable signal the system produces.

If our analysts stop leading, it is not a Human-Led AI SOC anymore, whatever the slide says.

Less drama, more detection

Sentry is not built to replace analysts.

It is built to give each of them a faster, better evidenced first draft, and to make sure the decision stays theirs.

That is what Human-Led AI SOC means to us in practice.

The AI does the gathering, the cross checking and the drafting at machine speed.

A person decides, and their name is on the verdict.

If you want to see how your alerts would be investigated, talk to our team.

Bring a hard case.

FAQ

Frequently asked questions.

  • Sentry Workbench is the investigation console QMasters designed to help SOC analysts turn case evidence into a reviewable first draft. AI gathers and cross-checks evidence, builds a timeline and proposes findings with citations. The analyst directs the investigation, accepts, edits or rejects each section, and signs the final verdict. It supports the analyst's judgment rather than replacing it. Sentry is designed and prototyped, with development beginning; the published screenshots use fictional demonstration data.

ABOUT THE AUTHOR

Gregori Nazarovsky
Cofounder & CTO, QMasters

Practitioners from the QMasters Security Operations Center. We run 24/7 monitoring, detection engineering, and incident response for organisations across regulated industries — and write here from the offense and defense work in front of us.

READY TO TUNE YOUR SIEM?

Tighter detections, fewer false positives.

Book a working session with a senior detection engineer. Bring a sample of your noisiest alerts and we will rebuild the rule with you.

Explore Managed Detection & Response

F-003 · CONSULTATION

Book 30 minutes. No slides.

A real working session with a SOC engineer — bring your alerts.