Security Context: Onboarding Agents to Your AppSec Team

Security Context: Onboarding Agents to Your AppSec Team

By

Mairtin O'Sullivan

Security Engineer at Synthesia

Create AI videos with 240+ avatars in 160+ languages

For anyone in the Product Security or AppSec world, the last year has revealed a repeating pattern as teams experiment with AI agents for security review.

Think About Your Last New Hire

When a strong senior security engineer joins your team, they aren't immediately productive. They need time to onboard.
Sure, they know how to read code. They understand how to spot vulnerabilities and have a solid handle on architectural best practices. All the fundamentals are there.

But for the first few weeks, and in some orgs even months, they still need help.

They're surprised by some legacy decision that seems bonkers, but makes sense for your org.

They're worried about specific threat actors or risks that were hugely important to their last company, but not yours.

They don’t yet know which services interact, which are crown jewels, and which are disposable.

Sound familiar? They don’t need to be smarter. They need context. For a human new hire, you've no doubt already done this:

If we can’t expect human new hires to have this context, why do we assume AI agents will magically just "get it"?

What Do We Mean by Context?

When we started testing different AI agents, we all agreed that context was key. But we also weren’t entirely entirely sure what we meant by that.

Framing it like onboarding a human helped us quickly align on a few key areas.

1. Architectural Context

How does your system fit together? Is it one giant monolith, or are you fully committed to micro-services? Which surfaces are external-facing? Where is the trust boundary between services?
An AI agent reviewing a PR in isolation has no idea whether it's reviewing a public API endpoint exposed to the Internet, or a highly trusted internal service with service-to-service auth in place.

2. Code Patterns and Conventions

Ask any developer and if they're not embarrassed by at least one key area of the codebase they support, they've probably shipped too late.

Every codebase has stuff that's weird, things that are legacy, things that made sense at the time but now seem bonkers. Are all of those best practice from a security perspective? Unlikely.

But you've reviewed them before, they're known safe, and if a new feature uses them, it's likely going to be grand. But if a new feature doesn't follow them, that's worth an AI agent flagging.

3. Threat Models

"Your threat model isn't our threat model": often stated, but without this context, how would an AI agent know?

Threat models document your reasoning: the realistic attack scenarios for specific features or services, and, crucially, what you concluded the actual risk was.

4. Risk Appetite and Accepted Risks

Risk appetite varies by org. Are you handling highly sensitive or regulated customer data where a breach would be a death knell for the business? Are you processing billions of time-sensitive transactions and a likely DDoS target?

You need to tell an agent what you actively care about. Making your priorities explicit is what separates findings that are actionable from findings that are merely technically correct.

The Problems

Agents All The Way Down

If you're like us when we started, you've probably already tried to give AI agents some of the above types of context already.
Maybe you've pulled some docs together that describe your architecture and used a RAG approach to give an LLM the context it needs? Maybe you've codified some of your risk appetite into prompts you've created? Maybe you've even written detailed agentic skills for specific kinds of security review?

Yup, we did all of that too, and if we had only one or two agentic tools or workflows, that might have been fine. However, we wanted multiple agentic workflows with us in the loop, covering document review, threat modelling and PR review, and ranging from fully automated workflows to agent-assisted manual review.

And that's when things started falling apart. Every agent had a slightly different version of reality. One had an updated threat model. Another was still working from an architecture overview that predated a significant refactor. When a risk decision changed, you had to remember everywhere it had been referenced and update each one manually.

In practice, you didn’t. What started as cool examples of what could be done became less valuable and far too much maintenance.

So we knew we needed to centralise our security context and make it consumable by any agent we were running, wherever and however we ran it. That's not a technical insight. It's basic information hygiene, but it took hands-on experience to appreciate how much it mattered at scale.

Context Has A Cost

Then there's the context window challenge. Yes, you could dump all this context into a single file or a couple of large files and load it into the agent's context window on every run, but that is going to make you sad quickly:

So the challenge was not just centralising our four context types. It was ensuring we could deliver the right slice of context to the right agent at the right time, without blowing the budget or burying the signal.

How We Approached It

We knew we wanted to give the AI agent right-sized, right-scoped context for the task at hand, and we wanted this to be centrally accessible to any agentic workflow we run.

So, perhaps unsurprisingly, we chose to create this context in a dedicated GitHub repo called (rather unimaginatively) appsec-security-context.

Why GitHub?

You likely already have wikis and tools for docs and runbooks, and many have APIs or MCP servers. We could have used those, but GitHub was the clear winner for us:

What’s In The Repo?

Thinking back to the four areas we outlined earlier that define context. We wanted a way to codify all this, and we landed on two core context types:

Codebase Context

In our previous post on PR scanning, we described why security context should be kept separate from codegen context. They pull in opposite directions: codegen context encourages following existing patterns, while security review often needs you to challenge them.

Codebase context captures the security expectations for a specific service. Each service/repo gets its own folder under codebase/, and the context is written in layers that match the repo structure, so an AI agent only loads what’s relevant to the code it’s reviewing.

Threat Model Context

We did not cover this in the previous blog post, and it is always a contentious topic. What is a threat model? The simplest answer is that it’s whatever helps you reason about risk and evaluate whether you have built the right controls.

For us, the key thing threat model context needed to solve was giving agents access to the reasoning, not just the conclusions. Not just "this risk is accepted", but why, under what assumptions, and what would need to change for that decision to be revisited.

But here's where it gets interesting. A threat model written for a human and a threat model written for an agent aren't the same thing, and trying to make one serve both purposes is where most threat modelling approaches fall down.

The human version is structured around reasoning, judgement, and readability. The agent version is structured for consumption, not comprehension. Same source of truth, fundamentally different shape.

So we create two, with both files live in the repo under threat-models/, one pair per feature.

What Workflows Use This Context?

We're currently using this security context across three distinct agents, and the difference it makes is visible in all. We'll cover each briefly here and plan to go deeper on each follow-up posts.

Doc Reviewer

What it does

How centralised context helps

PR Scanning

What it does

How centralised context helps

Backstop PR Scanning

What it does

How centralised context helps

Where We're Heading

Right now, as noted context is largely hand-authored. That works, but it's the most obvious bottleneck as the number of services and agents grows.

The next step is closing the loop: using the agent's own findings to prompt updates to the context.

The goal is a security context layer that improves over time and largely maintains itself, rather than a static set of docs that gradually becomes wrong and eventually misleading.

Build Your Own

Getting agents to do valuable security work isn't primarily a model problem, it's a context problem. Nothing we've built here is technically difficult. The value is in the thinking, not the tooling.

A few principles that have held across everything we've built: