About this role
Collinear AI builds environments, tasks, and evaluations that help frontier AI models improve at real work. We are growing our cybersecurity team and looking for someone who can turn practical security problems into environments where AI agents can investigate, act, and learn.
We recently released CWE-bench , a defensive cybersecurity benchmark with 100 held-out audit-and-patch tasks across 54 weakness types. Agents must find and fix vulnerabilities in real codebases, with checks that confirm the vulnerability is resolved and existing functionality still works. You will help build what comes next: richer cybersecurity environments, realistic tasks, and reliable ways to measure whether agents succeed.
You will own work from the initial security scenario through the runnable environment, task instructions, reference solution, and verifier. We are looking for hands-on security knowledge, strong programming skills, and curiosity about how AI agents fail.
Deep cybersecurity expertise is enough to get started—you do not need prior AI or machine-learning experience. If you know how to investigate vulnerabilities, reason about security failures, and verify that a fix works, we want to hear from you. We will teach you our AI tooling and evaluation workflows.
What you will do
- Build cybersecurity environments. Create isolated, reproducible environments with real repositories, applications, services, logs, and access controls. Make them straightforward to launch, reset, and evaluate at scale.
- Turn security work into tasks. Design scenarios around vulnerability discovery and remediation, application and API security, authentication and authorization, incident investigation, and system hardening. Define what the agent knows, which tools it can use, and what it must accomplish.
- Develop reference solutions and verifiers. Reproduce the underlying issue, implement a valid solution, and write checks that distinguish a real fix from a superficial workaround. Verify that attacks fail after remediation while legitimate behavior continues to work.
- Test the tests. Challenge graders with incomplete fixes, disabled features, hard-coded answers, and other shortcuts. Keep hidden solutions and test data out of the agent's environment, and separate genuine model failures from broken infrastructure.
- Run and analyze agents. Evaluate frontier models, inspect their tool calls and code changes, and explain where their security reasoning or execution breaks down. Use those findings to improve task coverage and difficulty.
- Contribute to benchmarks and research. Work with researchers and engineers to turn strong environments into evaluation suites and training data. Review other contributors' tasks and help document results for future releases.
Who we are looking for
- Practical depth in at least one area of cybersecurity, such as application security, vulnerability research, penetration testing, systems security, cloud security, or incident response. Evidence can come from internships, research, open-source work, bug bounties, CTFs, or independent projects.
- The ability to read unfamiliar code, reproduce a security issue, understand its root cause, and implement or assess a fix.
- Strong programming skills in Python and at least one language used in the systems you investigate, such as C/C++, Go, Java, JavaScript/TypeScript, or Rust.
- Comfort with Linux, Git, containers, debugging, and automated testing.
- Clear technical writing and careful judgment about what an evaluation does and does not demonstrate. You can explain why a task is realistic and why its grading is trustworthy.
- Curiosity about applying your cybersecurity expertise to AI, and a willingness to learn how to evaluate tool-using agents. Prior experience with LLMs is not expected.
Recent graduates and early-career engineers or researchers are encouraged to apply. A PhD, professional certification, or previous role at an AI lab is not required. We value demonstrated ability and the quality of your work.
Nice to have
- Disclosed vulnerabilities, accepted security patches, strong CTF results, or useful security tools and write-ups.
- Experience with fuzzing, static or dynamic analysis, reverse engineering, or building security labs.
- Familiarity with CWE and OWASP classifications and how they relate to concrete software failures.
- Experience building agent evaluations, adversarial tests, prompt-injection defenses, or reinforcement-learning environments.
*Examples of what you might build
- A repository audit where an agent must discover and repair an authorization flaw without being told where it is.
- A small service environment where an agent investigates suspicious activity from logs and configuration, then applies and verifies a remediation.
- An LLM application where an agent must repair a prompt-injection or tool-permission weakness while preserving legitimate functionality.
*When applying, include a project, code sample, security write-up, or research artifact that shows how you investigate a problem and establish that your solution works.