AI & DETECTION · 10 MIN
The Agentic SOC Is Running: A Five Minute Verdict
Our agentic SOC now reaches a documented verdict in about five minutes, and the point is that you can check it. Inside one run on real offenses.
Can an agentic AI SOC reach a verdict in five minutes, and can you trust it?
In our SOC the pipeline now reaches a documented verdict in around five minutes on many offenses, and the point is that you can check the work behind it. On 4 October a focused forty second follow up changed one case from ask_customer to need_tuning by reading two child events and citing them. The same system refused an unreliable baseline rather than turn a failed search into a first ever claim. It runs inside our SOC and delivers into an analyst queue. It is not released to customers yet. The automated quality grader is still switched off and the blind grade against our senior analysts has not happened, so five minutes is an objective we can show, not an accuracy guarantee.

Forty seconds.
Eight tool calls.
A recorded cost of $0.34.
That was one follow up inside a QMasters SOC AI investigation on 4 October.
Before it ran, the system could identify the account involved in an alert but could not explain what the account had done.
After it ran, the recorded findings described authentication activity, and the assessment became more specific: review the detection's watch list, and ask the customer whether this access was approved.
The verdict changed from ask_customer to need_tuning.
This is the post I have wanted to write since we started.
The pipeline is no longer a plan.
It runs inside our SOC, on real offenses, across many customer environments, and it reaches a documented verdict in about five minutes.
It is not released to customers yet, and I will be precise later about what that word means.
The reason I can show it now is not the speed on its own.
It is that a faster report only matters when we can explain why its conclusion follows from the work behind it.
*Customer names, account names, hostnames and network addresses are masked in this article.
Customers appear as Customer A through Customer F.
Diagnostic fields, recorded timings and decision labels keep their original meaning.
The examples come from our 4 October console records and internal investigation exports.*
The short version
What runs now.
A manager and worker pipeline investigates real offenses from our production SIEM across many customer environments.
The manager reasons, dispatches specialist workers for evidence collection, enrichment, ticket correlation and hunting, then writes the report.
Each worker returns one structured answer against a stated acceptance criterion.
The output lands in an internal queue that sits alongside StrongHold MCSS, where a human analyst owns the handoff.
How fast.
Completed runs on 4 October finished in roughly five minutes from queue to report.
Those durations include waiting on live SIEM and EDR queries, so they describe the investigation and not model speed.
Five minutes is an objective we can now show in the results, and the same screenshot shows cases that ran longer.
Why we built it ourselves.
A managed service spans many customer environments and many vendors, and an investigation standard has to travel across all of them.
We use existing models and the security tools our customers already run.
What we own is the orchestration: the objective each agent receives, the evidence it must return, and the way its findings become an analyst handoff.
That is the part no vendor ships for you, and it is the part that decides whether a fast answer is also a trustworthy one.
What one run shows.
A manager found that an alert named an administrator account but not its action.
It sent a forty second follow up with a testable finish line, received findings that described the action, and produced a more specific assessment.
It also rejected a baseline measurement it judged unreliable, rather than turn a failed search into a claim.
What we are not claiming.
The automated per run quality grader is still switched off, and every completed run shows grading skipped.
The blind grade against our senior analysts, the gate we set in Part I, has not happened.
The reports are internal CTI copies, so a recommendation does not prove a customer message was sent or that containment executed.
The investigation layer we are building
Our first post explained why we are building an operating layer across the security platforms we already use.
Our second post explained why execution records, not fluent summaries, have to govern any claim the system makes about its own work.
This post is the first one where I can point at a live result and walk the work that produced it.
QMasters SOC AI uses existing models and security tools.
We own the orchestration: the objectives agents receive, the evidence they must return, and the way their findings become an analyst handoff.
The architecture gives an investigation manager access to specialist work: event collection, threat enrichment, ticket correlation and broader hunting.
In the October example, the system collected the initial offense evidence first.
The manager then selected a child event follow up and a ticket correlation task, which ran together.
The useful question is how that selection works in a real case.
The operating model. A manager delegates specialist work and assembles the findings into an analyst handoff. We built this across the tools our SOC already uses.
An alert named the account, but left the action unknown
The case belongs to Customer A.
The detection monitored sensitive administrator accounts for any activity, and it involved a workstation, a management host and a domain controller.
For this article I will call them Workstation-01, Mgmt-01 and DC-01.
The account was Tier7Admin.
The detection description named a different administrator account and a different domain from the account in the alert.
That discrepancy raised a tuning question.
It did not establish that the observed access was legitimate.
The missing fact was the action itself.
An authentication request involving a management host leads to one line of investigation.
A group membership change or a privilege change on a domain controller could lead to a different severity and a different response.
The broad detection title could not settle that distinction.
The first pass recorded the gap plainly: the account's action, its host and privilege context, and its seven day baseline had not been examined.
Its verdict was ask_customer, with a recorded confidence of 0.4.
That was an incomplete investigation.
It was also a useful starting point, because it named the uncertainty the next step had to address.
The manager asked a question with a finish line
The manager's follow up targeted two specific Falcon child events.
It asked what each event showed: authentication, a group or privilege change, or another action.
The acceptance criterion required both child events to be read, with event type, action and host cited.
If the query failed, the worker had to return the exact error instead.
The manager also requested a QRadar count for the exact account name on the named hosts over the seven days before the alert.
That result needed its query, its time window and its distinct hosts, or a stated search failure.
This is what an evidence requirement looks like in practice.
The manager specified what the answer had to contain so that the next assessment could be checked.
A fluent paragraph about suspicious administrator activity would not have closed the question.
The follow up took forty seconds, used eight tool calls and recorded a cost of $0.34.
That cost belongs to this step.
It is not the all in cost of the investigation or of the managed service.
What came back, and how it changed the assessment
The recorded findings described two authentication actions.
| Child event | Recorded finding | Relevant hosts |
|---|---|---|
| First child | Successful Remote Desktop service request using NTLMv2. The worker excerpt names an Active Directory service access request, a REMOTE_DESKTOP type and a TERMSRV target. | The worker identifies Mgmt-01 as the target. The manager identifies Workstation-01 as the source. |
| Second child | Successful Kerberos CIFS ticket for the same account, as described in the manager's synthesis. | Mgmt-01 to the file service on DC-01. |
Neither reported child event described a group or privilege change.
The export available for this article contains a truncated worker finding and the manager's synthesis, rather than the complete raw results for both child events.
That limits independent reproduction from the article's source material.
It also means I should describe these as recorded investigation findings.
A Remote Desktop service request alone does not prove an interactive desktop session, and the supplied record does not give a Windows logon type.
With the action question addressed, the watch list mismatch became more relevant to the final recommendation.
The report returned need_tuning, with medium severity and medium confidence.
The customer question became precise:
Was this account's access to the management host approved, and should the watch list include it?
The system had narrowed the technical uncertainty.
The customer still held the authorization fact.
| Decision input | First pass | After the follow up |
|---|---|---|
| Account action | Unknown | Reported RDP service request and Kerberos CIFS authentication |
| Group or privilege change | Unresolved | Neither reported event describes one |
| Rule scope | Description and observed account disagree | The discrepancy remains relevant to tuning |
| Verdict | ask_customer | need_tuning, medium severity |
| Next action | General authorization gap | Watch list review and a specific access approval question |
This is the evidence for the method's contribution in this case.
Additional work produced findings that made the assessment and the handoff more specific.
It is a comparison inside one case, not an accuracy benchmark.
The recorded decision before and after the forty second follow up. The account action moved from unknown to two cited authentication events, and the handoff became a specific approval question.
The zero we could not trust
The follow up returned with status partial.
The action question had an answer.
The seven day baseline did not have a reliable one.
The manager identified two problems with the QRadar zero: the search window was off by twenty minutes, and the account name field was unreliable.
It explicitly declined to treat that zero as an account baseline.
Without that distinction, a report could have turned an unsuccessful measurement into a claim that the account had appeared for the first time.
Even a valid seven day zero would only describe that measured window.
It would not establish first ever.
The opposite claim would have been equally unsupported.
We could not call the access routine without a valid historical measurement.
I consider the decision to reject that result part of the value on display here.
The manager used the findings it judged useful and preserved the failed baseline as a limit.
A partial result still contributed to a better assessment without becoming a claim of complete coverage.
Rule recurrence and account history are different measurements
Ticket correlation found 27 tickets for the rule over 30 days.
The returned list contained 25 rows because the search result was capped.
The count, rather than the displayed rows, supplied the recurrence total.
That is a small implementation detail with a real consequence.
Counting only the returned rows would understate how often the rule had fired.
The entity searches found no matching ticket text for the account or for the two named hosts.
Those results answered a ticket history question.
They did not count the account's activity in the SIEM.
The final assessment therefore had to keep several observations apart.
- The rule had substantial ticket recurrence.
- The account and named hosts had no matching tickets in the measured searches.
- The exact account event baseline remained unreliable.
Twenty seven rule tickets can support a discussion about rule noise.
They cannot establish that this particular access was benign.
Zero matching tickets cannot establish that the account never used those hosts.
The scope of a measurement has to survive into the conclusion.
The follow up fit inside work that was already running
At 19:32:15Z, the child event follow up and the ticket correlation started together.
The follow up finished after forty seconds.
Ticket correlation continued for 1 minute 54 seconds.
The manager then reassessed the findings and produced the final report.
| Stage | Start, UTC | Recorded duration |
|---|---|---|
| Initial evidence review | 19:31:09 | 18.2 seconds |
| Initial synthesis | 19:31:27 | 23.5 seconds |
| Manager dispatch | 19:31:50 | 24.1 seconds |
| Targeted child event follow up | 19:32:15 | 40.0 seconds |
| Ticket correlation, alongside the follow up | 19:32:15 | 1 minute 54 seconds |
| Manager reassessment | 19:34:10 | 14.6 seconds |
| Final report | 19:34:24 | 1 minute 24 seconds |
The detailed timeline ended at 19:35:49Z.
That gives 4 minutes 40 seconds from the first agent's start to completion.
The run entered the queue at 19:30:48Z, which makes the elapsed time from queue to completion about 5 minutes 1 second.
Those are different measurements.
For a service promise, queue time is the one that matters.
The overlap shows how useful follow up work can fit inside an efficient investigation.
It does not tell us exactly how much faster this run was than a serial version.
We did not supply a controlled comparison for that claim.
The same console snapshot shows a separate completed case at 4 minutes 33 seconds, alongside completed cases at 5 minutes 2 seconds, 5 minutes 44 seconds and 9 minutes 15 seconds.
Our five minute objective is visible in the results.
The sample also shows why it stays an objective rather than a universal guarantee.
Completed runs reached the five minute range. The same snapshot shows longer cases, and every completed row shows grading skipped. This sample does not establish a service level.
Customer B: a raw log example, the VPN password spray
The administrator case shows the decision loop.
A FortiGate case for Customer B gives a more direct example of the event fields supporting an assessment.
The exported report includes this selected raw evidence, with the source address masked:
action="ssl-login-fail"
tunneltype="ssl-web"
remip=[source address redacted]
user="temp"
reason="sslvpn_login_unknown_user"
The report's evidence entries describe the same source trying different usernames.
| Time, UTC | Attempted username | Recorded result |
|---|---|---|
| 13:41:10 | temp | SSL VPN login failure, unknown user |
| 14:05:14 | cad | SSL VPN login failure, unknown user |
| 14:39:29 | paul | SSL VPN login failure, unknown user |
Eight of the offense's sixteen events were read.
All eight were failures of this kind.
The report classified the activity as a True Positive password spray.
It also reported 1,063 failures over a 14 day historical search, and about 18 attempted usernames.
The historical count adds scale, but its underlying query receipt is absent from the supplied export.
The raw excerpt and the report's event entries provide different levels of support, and we should label them accordingly.
The next action was to block the source and check for successful VPN sessions.
That final check matters.
Evidence of failed attempts supports an attack assessment.
It does not settle whether another attempt succeeded.
The report kept that uncertainty and made the response conditional on what the success check found.
This is what a useful investigation adds to an alert.
It explains the attempted behavior, states what the observed events establish, and gives the analyst a concrete next check.
Customer C: an endpoint example, one word can overstate containment
Another case, for Customer C, involved a test executable in a developer scratchpad.
The report described a process chain involving an AI coding assistant, PowerShell and a test build.
The raw endpoint disposition contained the following wording:
process was blocked from execution and quarantine was attempted.
The narrative described the file as quarantined.
Those statements carry different certainty.
The raw excerpt establishes blocking and an attempted quarantine.
It does not, by itself, establish that quarantine completed.
Development context also cannot establish approval or harmlessness.
The report's next step was to confirm the build's purpose before proposing a narrow exception.
I want this discrepancy in the build story, because it gives us a concrete quality problem to fix.
A report can assemble useful context and still overstate one outcome.
The evidence has to constrain the wording all the way through to the final paragraph.
Customers D, E and F: the same discipline across platforms
The other supplied cases show the same need for precise interpretation across different telemetry.
Customer D, a Windows off hours login.
The alert was a successful interactive login outside working hours, with logon type 7 and an elevated token of no.
The report read logon type 7 as a session unlock rather than a fresh sign in, which is the correct reading of that field.
That changed the question.
The remaining uncertainty was not what happened, it was whether the unlock was expected, so the verdict was Need Customer Approval.
Customer E, an Okta login from a flagged address.
A login arrived from an address that carried low confidence reputation signals, a low score on one reputation source and a single detection out of more than ninety scanners.
The investigation did not stop there.
It found a managed and registered device associated with the user, and the same user and address appearing across Okta and several business applications for weeks beforehand.
It also recognized that a hash like value in the record was an Okta device token field, not a file, and it did not report it as malware.
The result was a specific question rather than an alarm: confirm expected use before writing a tenant specific exception.
Customer F, a SentinelOne file detection.
The endpoint flagged a file as malicious and reported it mitigated.
The file was a known attack framework script found inside a repository clone on a developer's laptop.
A signature verdict of malicious is a real signal, and it still does not establish whether the tool ran or whether its presence was authorized.
The next step was to confirm tool authorization and review the execution evidence with Tier 2, not to close the case on the label alone.
The breadth is the point.
Each platform has its own event semantics, and the shared investigation method has to preserve them rather than flatten them into one confident sentence.
| Customer | Platform | Evidence that mattered | Resulting handoff |
|---|---|---|---|
| A | Falcon and QRadar | Reported action differs from a privilege change | Watch list review and an access approval question |
| B | FortiGate | One source, many attempted users, failed VPN access | Block the source and check for successful sessions |
| C | Falcon | Blocked, with quarantine attempted, in a development build | Confirm the build before a narrow exception |
| D | Windows | Logon type 7, read as an unlock | Expected use question for the unlock |
| E | Okta | Known device and weeks of application history | Tenant specific authorization question |
| F | SentinelOne | Malicious signature plus tool and repository context | Confirm authorization and escalate to Tier 2 |
Sources and evidence notes
The public build narrative began with Building the Agentic SOC: Why We Stopped Waiting, published on 7 August 2026, and continued with Agentic SOC Part II: The Agent's Word Is Not Evidence, published on 12 September 2026.
The operational examples here come from the supplied 4 October 2026 execution snapshot and internal report exports.
They are selected examples, not a representative benchmark.
Their source limits are stated alongside the claims they affect.
- Customer A, central example.
Internal console export, 4 October 2026, across the initial synthesis step, the manager follow up brief, the targeted follow up, the reassessment and the final report.
The follow up excerpt is truncated, and the complete raw child event results and a valid exact account baseline are not included in the available export.
- Timing.
Supplied run list and detailed step screenshots, 4 October 2026.
The detailed run spans 19:31:09Z to 19:35:49Z after entering the queue at 19:30:48Z, and about 5 minutes 1 second is calculated from these displayed timestamps.
Step durations are rounded, and concurrent durations must not be added to derive elapsed time.
The separate 4 minutes 33 seconds case is a different offense from the Customer C endpoint example.
- Customer B, FortiGate.
Internal report export.
The selected raw fields are from one event, with the source address redacted, and the three rows correspond to recorded evidence entries.
The export includes the reported 14 day count, but not its underlying query receipt.
- Customers C to F.
Internal Falcon, Windows, Okta and SentinelOne report exports.
The quoted endpoint wording is a substring of the recorded pattern disposition description.
Complete raw telemetry for all findings is not reproduced in those exports.
For more about QMasters and our managed security services, contact our team.
FAQ
Frequently asked questions.
On 4 October completed runs finished in roughly five minutes of wall clock from queue to report, with examples at 4 minutes 33 seconds, 5 minutes 2 seconds, 5 minutes 44 seconds and one at 9 minutes 15 seconds. Those durations include waiting on SIEM and EDR queries, so they measure the investigation, not model speed, and the sample is not a service level agreement.
Not yet. It runs inside our SOC and delivers investigations into an internal analyst queue. Humans own every customer facing word. We are moving it into live operations and preparing a wider release once it clears the gates we set in public, which include a blind grade against our senior analysts.
The first pass could name the administrator account in the alert but not its action, so it returned ask_customer with low confidence. A forty second follow up of eight tool calls read two child events, found a Remote Desktop service request and a Kerberos ticket, and the recorded verdict became need_tuning with a specific access approval question for the customer.
Only because the work is inspectable. You can read the question the manager asked, the evidence that came back, and the limits it preserved, including a baseline measurement it refused to trust. The automated grader is off and a blind grade against senior analysts has not run, so we present selected examples with their source limits, not an accuracy rate.
The reports are internal CTI copies. A recommendation to block an address or confirm a build does not prove a customer message was sent or that containment executed. Those remain human decisions confirmed against the relevant control.
ABOUT THE AUTHOR
Practitioners from the QMasters Security Operations Center. We run 24/7 monitoring, detection engineering, and incident response for organisations across regulated industries — and write here from the offense and defense work in front of us.