Brief Reflections on Security as Engineering, LLMs, and the Meaning of Research
Contents
-
Introduction: The Place of Computer Security Research in Computer Science
-
Main Discussion
-
An Unfunny Joke About LLMs
-
My Entirely Personal Distinction Between Papers That Use LLMs Well and Papers That Don't
-
But Why Should We Use LLMs Well?
-
Conclusion: A Brilliant, Clear-Cut Answer Spanning Philosophy, AI, and Cybersecurity
Introduction: The Place of Computer Security Research in Computer Science
I am a graduate student at a university in South Korea, hoping to earn a PhD in computer security. More precisely, I am still in the master's stage of that infamous arrangement known as an integrated master's and PhD program. Calling myself a PhD student feels a little dishonest; calling myself a master's student does not quite account for how much time I have left.
Our lab once hosted a seminar by someone who had earned their PhD at one of the world's leading security research labs. After the presentation, our professor promptly shoved our guest into our lab, where roughly five hours of debate and Seoul restaurant recommendations ensued under the pretense of practicing conversational English.
One topic that came up was our guest's disillusionment with the security research community. Putting their argument as gently as possible:
Computer security research is, fundamentally, not all that different from pseudoscience.
What they actually said was a little more radical. This post, however, is not an attempt to address that sweeping claim in its entirety. I want to borrow just one uncomfortable question from it: how we evaluate security research papers.
Their particular objection concerned newly discovered vulnerabilities.
First, it is strange to treat the discovery of previously unknown vulnerabilities as a requirement for demonstrating a paper's contribution. Second, it is also strange to use the number of such discoveries as the central evidence that a proposed technique is superior. Third, if researchers try numerous programs and configurations and then present only the successful cases, the entire exercise can become a vast operation in cherry-picking.
At the time, I strongly disagreed. If a new technique discovered new vulnerabilities in real programs, I thought that was fairly persuasive evidence of its practical usefulness.
We never came to agree, but we buried the argument at a convenient point so that we could exchange restaurant recommendations. (In truth, seniority won. I backed down.)
Looking back, I think the issue would have been more precisely framed as follows:
In security research, how much does the fact that something actually worked establish about its contribution?
When evaluating computer security research, I think it helps to consider both scientific explanation and engineering validation. At least the kind of systems security research I am interested in is often closer to designing solutions that work in complex, human-made systems than to discovering laws that already exist in nature.
In this kind of research, it is easier to assess a technique's value when we consider its effectiveness in real environments, failure conditions, cost, and reproducibility alongside the elegance of the method.
Discovering previously unknown vulnerabilities is one form of evidence that a technique works in the real world. (Or so I think.) It shows that the proposed technique can find meaningful problems in code at a realistic scale, rather than only in synthetic examples.
But we should distinguish between the fact that a technique found new vulnerabilities and the claim that the number of discoveries demonstrates its general superiority.
The number alone makes it difficult to judge whether the technique will also outperform alternatives on other programs and vulnerabilities. Unless we also know which programs were selected, how many targets were tested, and what characterized the targets on which the technique failed, it may remain unclear how much of the success came from the technique and how much from the choice of experimental targets.
In other words, there is quite a distance between “we found new vulnerabilities” and “therefore, this method is generally superior.”
That distance is what this rather long introduction is about.
While doing research recently, I found myself wondering:
If practical success matters in engineering research, how much does getting good results with an LLM justify the choice to use one?
An Unfunny Joke About LLMs
I have a favorite joke about LLM research.
When I kick a chair with my left foot, it goes thud.
When I kick it with my right foot, it goes clatter.
When I kick it with both feet, it goes clatter, thud.
A truly astonishing discovery.
If I carried out hundreds of chair-kicking experiments, quantitatively analyzed the differences in sound under the left-foot, right-foot, and both-feet conditions, and submitted a paper on these findings, I would probably have to worry about a desk rejection before anything else.
But replace the chair with an LLM, and things seem to change a little.
When I enter the left prompt, the LLM goes thud.
When I enter the right prompt, the LLM goes clatter.
When I enter both prompts, it goes clatter, thud.
Suddenly, a title like this becomes possible: “An Empirical Study of Composite Response Behavior in Large Language Models Under Varying Prompt Compositions.” With a bit of luck, it might even get into a top-tier conference.
Of course, I do not mean that conferences are actually this simplistic. The joke is closer to an elaborately constructed straw man designed to explain my frustration.
(Observing a new and complex object can certainly be valuable. What I wanted to express through this joke was a suspicion: are we a little more generous toward observations we would otherwise consider trivial, simply because the object happens to be an LLM?)
People around me were kind enough to sympathize with the joke. So I took it one step further:
When I read some security papers that use LLMs as magic hammers, I find myself wanting a little more explanation connecting the reason for using them to the results they produce.
This view was reinforced by my advisor's position that papers failing to adequately justify their use of LLMs should be rejected, and by my own experience of being repeatedly torn apart over the same issue in lab meetings.
But my advisor and our lab meetings are not scientific evidence. The number of times I have been torn apart in lab meetings might, at least, be statistically significant, but that alone does not support a general conclusion.
In the beginning, there was the problem definition.
At least, that was supposed to come first.
So I decided to come up with some criteria of my own.
My Entirely Personal Distinction Between Papers That Use LLMs Well and Papers That Don't
I, too, have research to do. And I have to use LLMs. Please do not ask why. That is where my paycheck comes from.
Making a living is a sufficiently strong motivation for research, but it is generally not accepted as a methodological justification. I therefore needed to figure out what distinguished research that, in my view, used LLMs well from research that did not.
To do so, I conducted a rigorous, fair, and reproducible paper selection procedure.
(Inside these parentheses are the dataset selection procedures you would expect from a magnificent survey paper, along with proof that the process was fair and impartial. If you find this difficult to accept, picture the sheep inside the box.
I trust that you are now satisfied.)
To prevent international incidents, I will not disclose the names of the papers.
The studies I liked generally offered reasonably good answers to three questions.
First, what does the LLM do, and what advantage does it offer?
In studies I found persuasive, the problem assigned to the LLM was defined relatively concretely.
Examples include connecting security requirements written in natural language to code, identifying violations of similar security invariants in syntactically different code, or generating candidates needed for analysis from incomplete context.
An explanation like “we used an LLM because semantic understanding is required” did not, by itself, give me enough to assess whether the design was appropriate.
The choice becomes easier to understand when the paper makes clear which semantic relationships it addresses, what goes into and comes out of the LLM, how the correctness of its outputs is judged, and what role it plays in the overall system. Comparisons with simpler rules, retrieval, existing program analysis techniques, or smaller models also help establish what advantage the LLM provides.
An LLM does not have to be the only possible solution. Better performance, broader applicability, or lower development costs can all be perfectly good reasons to choose one. Comparing these benefits with inference costs and the burdens of reproducibility and opacity would help us assess how substantial the advantages are.
If the decision to use an LLM came first, though, I become especially curious about how that choice was compared with alternatives. Trying a new tool can lead to the discovery of a suitable problem, so the order in which the work began is not enough to judge the research.
What I found less satisfying were cases where the problem seemed amenable to rules or retrieval, yet I could find little justification for choosing an LLM beyond a slight improvement in the final performance numbers. Without enough explanation connecting the problem to the results, it is difficult for a reader to tell whether the LLM is a component the design calls for or an ornament that makes the paper look more impressive.
Second, where did the performance improvement come from?
I work, therefore I contribute.
Descartes never said this. But some evaluation sections leave me with a similar impression.
If adding an LLM improved the results, I also want to know where that improvement came from.
Was it the prompt? The retrieved context? The larger model? Or were external tools doing the actual judging?
Ablation studies and baseline comparisons help us examine the contribution of each component. Adding every component at once and reporting only the final number is like changing the left foot, the right foot, the chair, the floor, the recording equipment, and the humidity in the room, then claiming that the sound changed. We may be able to confirm a difference in the results without being able to tell what caused it.
Third, when does it get things wrong, and what breaks when it does?
All successful demos are alike; each unsuccessful system fails in its own way.
We need to examine whether judgments become unstable as inputs or context change, and how sensitive the results are to changes in the model or prompt. If a modest increase in accuracy comes at the cost of a several-dozen-fold increase in analysis costs, considering that trade-off also helps us judge what the result means.
The consequences of errors depend on the role assigned to the LLM. If a person reviews the suggested vulnerability candidates, false positives can be filtered out. But if analysis targets are removed because the LLM judges that “this path is safe,” a single wrong answer can cause a vulnerability to be missed.
Not every LLM output needs to be checked through formal verification. What matters, I think, is choosing a level of validation appropriate to the impact an error could have on the research conclusions and on actual security.
To summarize, research that uses LLMs well, in my view, generally looks like this:
Research that makes the LLM's role and the reason for choosing it clear, allows us to examine each component's contribution to the performance improvement, and explains both the conditions under which errors occur and their consequences.
These criteria are subjective.
But Why Should We Use LLMs Well?
And so I came to hold two views.
First, it matters that security research works in real systems.
Second, considering why an LLM was chosen and the limits within which it can be trusted makes its results easier to interpret.
Then I began to wonder whether I was applying stricter standards to LLMs than to other tools.
Is this a double standard?
When evaluating the choice of a tool in a security paper, we usually consider how well it fits the problem. We might ask why a particular search strategy, analysis technique, or solver was chosen, but that is a far cry from demanding philosophical proof that it is the only possible solution.
Yet are we especially prone to asking papers that use LLMs, “Is an LLM really necessary?”
Isn't building a working system enough?
If it finds more new vulnerabilities, outperforms existing methods, and solves real problems, doesn't that establish a contribution?
This is not an objection I can easily dismiss.
My current, tentative answer is:
It seems reasonable to apply the same evaluation principles to LLMs as to other tools. However, when some factors are difficult to hold fixed and control experimentally, additional explanation may be needed about the reasons for the choice and the meaning of the results.
This is not because I think LLMs are particularly mystical or evil.
The results of a system incorporating an LLM can depend on many factors: the influence of training data, the supplied context, generated code, and the behavior of external tools, among others. It may therefore be difficult to explain an observed success in terms of a single mechanism.
In particular, if the system uses a commercial API whose model version is difficult to pin down or whose change history is difficult to check, we need to consider how much control the experiment actually had over the variable called “the model.”
When information about the model is not sufficiently public, it can be difficult to determine whether the internal version or policies changed under the same name, or whether the evaluation targets were included in the training data. How consistently the model produces results for the same input is another question that needs to be examined separately.
This uncertainty about reproducibility can affect more than the effort needed to obtain the results again. It can also affect our ability to explain what, exactly, we experimented on.
Under these conditions, performance numbers can admit more than one interpretation. Final scores alone make it difficult to distinguish whether an improvement came from the model's semantic reasoning, retrieval modules and external tools, or its ability to pick up on superficial features of the evaluation data.
The reason I expect additional explanation is that I want to know how much these uncertainties affect the conclusions of the research.
Of course, we can ask similar questions of other tools. Explaining the failure conditions of heuristics, the boundaries of soundness and precision in static analysis, or the influence of benchmarks and seeds in fuzzing research also helps us evaluate the results.
For LLMs, too, it seems more reasonable to determine the level of explanation required by what the experiment could control and what uncertainty remains in its results, rather than by the name of the tool.
Viewed this way, perhaps our guest's argument and mine were not entirely opposed after all.
Discovering new vulnerabilities and improving performance with LLMs are both valuable achievements. But additional evidence may be needed before we can infer a technique's general superiority from the number of discoveries, or conclude that an LLM has addressed the essence of a problem simply because a number went up. Distinguishing the observed achievement from the scope of our interpretation makes it easier for readers to decide how much they can accept.
A tool's success produces results. Explaining the conditions under which those results arose, how far they can be generalized, and where the method fails allows us to assess the research contribution more clearly.
Conclusion: A Brilliant, Clear-Cut Answer Spanning Philosophy, AI, and Cybersecurity
Philosophers have long argued over how much an observation can justify a particular explanation.
If a single observation is compatible with several different explanations, observing a successful result does not, by itself, establish any one of them.
The fact that a shadow on a cave wall achieves 93.7% accuracy does not mean we fully understand the object outside the cave.
Security research has its own circumstances that make these judgments difficult. Attackers and defenders can adapt their approaches in response to one another, and systems and models keep changing. Unsuccessful attacks and undiscovered vulnerabilities can also be difficult to observe.
LLMs can be powerful and convenient tools in this environment.
Using an LLM as a hammer can be useful, too. But before we move from “something broke” to “we drove the nail in properly,” I would like to hear a little more explanation of what happened in between.
So:
In engineering research, the fact that a tool worked has value in itself.
Explaining when, why, and how far we can trust it helps us better evaluate and use that achievement.
Of course, I have discovered a far more brilliant and clear-cut answer encompassing philosophy, AI, and cybersecurity. But the margins of this blog are too narrow to contain it.