AI Narrowed It to 100 — The Rest Is on Us
There's no denying that AI has dramatically accelerated the speed of forensic analysis. Not long ago, analysts had to manually sift through an overwhelming volume of logs, line by line. Now, with the help of AI, a significant portion of that workload has been reduced. From artifact parsing to timeline reconstruction to flagging suspicious items, the scope of what AI can handle has clearly expanded.
But there's something I've been feeling more and more strongly through recent casework. We must not become too absorbed in AI's analysis results. AI is undeniably fast and convenient. But fast and convenient results don't necessarily mean accurate results.
The Line AI Still Can't Cross
AI already does a lot of things well — extracting patterns from massive log sets, identifying traces that resemble known attack techniques, and pointing out items an analyst might overlook.
But there's one area where I feel AI still falls short. That's distinguishing the thin line between normal and abnormal.
In incident response, the traces that matter most are often not the ones that are clearly malicious, but the ones that look normal yet are abnormal within context. Why is a legitimate process running at that particular hour suspicious? Why does a login record that looks identical to any other raise a red flag? These are judgments that can only be made with a full understanding of the incident's context.
AI doesn't seem to handle this gray area well yet. It catches known patterns effectively, but it has limitations when it comes to judging subtle differences that only become meaningful within a specific context.
That's why I believe what AI provides will always be, at best, "elements worth investigating." Whether those elements are actual traces of an attack or simply part of normal operations — that final call is ultimately on the analyst.
One Image In, One Report Out
These days, AI-powered agentic services are becoming increasingly common. You can feed an entire forensic image as input, sit back with a cup of coffee, and just keep hitting the Allow button. Before you know it, you have artifact parsing, timeline construction, suspicious item analysis, and even a preliminary report — all done for you.
It's incredibly convenient. And most of the time, the content in that report is actually correct.
But can we really say the analysis is complete based on that report alone? If we never opened the actual tools to verify anything ourselves, and just signed off on the AI-generated report, I don't think we can call that analysis.
The fact that most of it is correct actually makes the problem harder. When nine out of ten findings are right, it's easy to assume the last one is too. But in practice, I've had more than a few experiences where a handful of misanalyses threw the entire investigation off track.
If those few errors go uncaught, they lead to wrong conclusions. And those conclusions become the basis for client reports, response strategies, and sometimes even legal proceedings. That's why blindly trusting AI-generated results is dangerous.
Reducing 10,000 to 100
That said, I'm not trying to argue that AI is useless. Quite the opposite.
Where an analyst previously had to manually review 10,000 items one by one, AI narrows that down to 100. It filters out suspicious items, prioritizes them, and focuses the analyst's attention where it matters most. That alone, I believe, is more than enough to justify AI's value.
But we shouldn't assume all 100 of those items are correct. We have no way of knowing how many are accurate and how many are wrong. So ultimately, all 100 need to be fact-checked. What AI reduced is the scope of what needs to be reviewed — not the act of reviewing itself.
If we lose sight of this, AI becomes more of a liability than an asset. The moment we start thinking "AI handled it, so I don't need to look," the accuracy of our analysis becomes entirely dependent on the accuracy of AI.
What I really want to say is this: AI is not always right, and we need to continuously build the foundational knowledge required to verify whether AI's results are actually correct.
When AI flags a specific artifact as evidence of an attack, you need to understand the structure and meaning of that artifact to confirm whether the judgment holds. When AI constructs a timeline, you need to know the context behind each event to catch missing pieces or incorrect connections.
The attitude of "AI says it's right, so it must be" is no different from voluntarily stepping down from the analyst's role.
And this isn't a story limited to forensics. In any field, leveraging AI requires having the domain knowledge first. You need the eyes to judge whether AI's answer is right or wrong before you can truly use AI as a tool.
Love Att&ck
There's something else I've been feeling strongly while working incident response cases recently. To make the judgment that "this is correct," and to decide which areas need further investigation, understanding the tools that real red teams use and how they work is incredibly important.
Attackers do everything in their power to leave no trace — wiping logs, executing tools only in memory, masquerading as legitimate processes, tampering with timestamps. You need to understand these techniques to even begin thinking about what traces might remain despite all those efforts.
If a specific technique was used, you can determine which logs would be generated and which wouldn't. Without this understanding, an analyst only sees what's right in front of them. Analysis gains depth not just from examining what's present, but from thinking about why what's absent is absent.
Finding the Missing Pieces
In CTF competitions, there's always a guaranteed answer to find. But in real incident response, the data available for analysis is limited. Most logs are either partially lost or intentionally deleted by attackers before the investigation even begins. Having limited visibility is more the norm than the exception.
AI draws its conclusions from the data it's given. But if that data itself is incomplete, then AI's conclusions are inevitably incomplete as well. Initial access vectors might be missing due to lost logs. Parts of the attacker's activity might be entirely invisible because the traces were wiped.
What Doesn't Change in the Age of AI
Through all these experiences, something has become clear to me. To avoid misjudgments born from blind faith in AI, we need to steadily develop the foundational knowledge required to assess the integrity of our analysis results. And I've come to realize that this means broadening our interest beyond just web exploitation or malware analysis — it means understanding the diverse tools and techniques used in a single attack chain and the principles behind them.
Understanding what tools attackers use, how they cover their tracks, and what paths they take — I believe this is not optional but essential for a forensic analyst. Only then can we look at AI's output and distinguish: "this part is likely correct," "this part needs additional verification," and "this part is beyond what AI can reliably judge."
AI makes analysts faster. But it doesn't make analysts more accurate. Accuracy still comes from human knowledge, experience, and judgment.
Be grateful to the AI that narrowed 10,000 down to 100. But don't hand over the job of verifying those 100 to AI as well. In the end, the final line of analysis must be written by a human.