Skip to main content
PrismCV
JobsExtensionPricing
LoginCheck Your Resume
Check Your Resume
← Back to all jobs

Software Engineer - Offensive Security

GhostEye · New York, NY, US

$180k - $250k
On-site
Full-time
Mid
Check your resume against this jobApply on Ycwaas

Job Description

About GhostEye

Most security programs tell companies whether they appear compliant. GhostEye tells them whether they can withstand a real attack. We're building the always-on red team for modern enterprises. GhostEye emulates the attacks adversaries use today, including help-desk vishing, deepfaked executives, MFA fatigue, and the technical payloads that follow. We identify what is actually exploitable, help customers close the gaps, and retest until the fixes hold.

Our founder previously led red-team work at BlackRock and conducted offensive cyber operations at MITRE. We're bringing that experience to one of cybersecurity's hardest problems: helping organizations continuously test their defenses against the attacks they actually face.

About the role

Attackers already use synthetic voice and video, and the volume is climbing fast. They call help desks and sound like your CFO. They join video calls as an executive who never joined. Most companies have no idea whether their people hold up against this. The only honest way to find out is to run it against them.

That's what you'll build. You'll own the AI behind our social engineering simulations: voice agents that can hold a live help-desk call, synthetic video for executive impersonation, and pretexting grounded in real research about the target company. All of it under authorization, with controls that hold up to scrutiny.

This is one of the most important roles on our engineering team. You'll report to our Head of Engineering and work closely with our security researchers. You'll also be the first person here whose full-time job is applied AI, so you'll set the technical direction for it.

One thing worth saying plainly. Most of your work is the AI above, and that's why the job exists. You'll also build APIs, write the service that runs a campaign, and sometimes ship the screen a customer looks at afterward. If you want to work only on models, this is the wrong role.

What you'll own

Real-time voice. The pipeline that runs a live help-desk call: streaming speech recognition, voice cloning, dialogue management, barge-in and turn-taking, and telephony. All of it inside a conversational latency budget. A half-second of dead air ends the call, which means you're measuring your own lag instead of the customer's controls. Most of this is a streaming systems problem. The model choices are the easy part.

Synthetic video. Face and voice quality good enough to test whether a company's identity checks hold up on a live video call. Demo quality isn't the bar. A trained employee following procedure should not be able to tell.

Pretexting. Systems that take open-source research on a target company and produce pretexts, phishing content, and conversational strategy that match what a real operator would do. The hard part is grounding. An agent that invents details gets caught, and one that invents details about real people creates liability.

Evaluation. Working out whether a simulation was realistic, whether it worked, and why. No benchmark exists for this, so you'll invent most of how we measure it. That measurement is a large part of what customers buy. Telling someone they failed is worthless without telling them exactly where.

Data and labeling. Every authorized engagement produces data that exists nowhere else: real employees at real companies responding to synthetic attacks, with ground truth about what happened. Capturing it well is your job.

  • Turn-level annotation, not just call outcomes. Which turn raised suspicion. Which part of the pretext carried it. Where the audio broke down.
  • Getting our researchers' judgment out of their heads and into labels, preference data, and eval sets. It's the most durable thing you can build here.
  • Annotation tooling good enough for senior red teamers, whose time is the real constraint.
  • Label quality as engineering work: agreement, adjudication, and drift as attacker tradecraft changes.
  • Grounding labels, so we can measure how often the model invents a detail that would get an operator caught.
  • Refusal datasets, kept as a regression suite so guardrails don't erode with the next model swap.
  • Privacy throughout. This corpus holds real people's voices and mistakes, under the same custody and destruction rules as everything else. Build for that from the start.

Safety, authorization, and containment. A large part of this job, and engineering work rather than paperwork.

  • Authorization enforced where the action happens. Scope, target lists, numbers, domains, and time windows bound to runtime and checked at the moment of the call, not at config time. An out-of-scope call should be impossible to place. A customer should be able to stop an engagement mid-call.
  • Likeness consent. Cloning an executive's voice or face needs their recorded consent, tied to one engagement, time-boxed, revocable. Provenance for every artifact we generate.
  • Artifact lifecycle. Cloned voices and generated media are the most dangerous things we hold. Encryption, tight access, no egress, hard tenant isolation, and provable destruction when an engagement closes.
  • Watermarking, so media we generate can be identified later. It protects customers, protects us, and is increasingly the law.
  • Autonomy limits. What an agent may do without a human, and what it must never do: no improvising past the approved pretext, no coercion, no targets outside scope, nothing reaching personal devices or family. We record that a credential was offered, never the credential.
  • Target welfare. The people on the other end are our customers' employees. Stop conditions, engagement boundaries, and data handling that keeps results from being used to punish individuals.
  • The legal constraints that shape the architecture: call recording consent by state, automated calling rules, biometric and deepfake statutes, GDPR for employees outside the US.
  • Audit trails good enough to reconstruct an engagement for a customer's legal team a year later.
  • Attacking our own guardrails. We're a red team. These controls should be under constant attack from us, with tests that prove the system still refuses what it should.

Technical direction. What we build and what we buy, in a market that changes every month, plus the research roadmap alongside our security team.

The product around it. Services, APIs, orchestration, and the customer-facing pieces that turn a simulation into something a CISO can act on, all in TypeScript. Whoever builds the capability here also ships it.

What we're looking for

  • You've shipped applied AI to production, not just prototyped it. Real users, real latency budgets, real failure modes.
  • Depth in one of: real-time speech (recognition, synthesis, voice conversion, streaming audio), generative video or face synthesis, or LLM engineering with retrieval, tool use, and evaluation.
  • You're good at latency. Real-time pipelines, WebRTC, or low-latency inference count for more here than model training.
  • Strong general engineering. We build in TypeScript, and you'll reach for Python where the ML tooling requires it. We care that you can build and run real services, not which language you did it in, but you should be good at TypeScript or ready to get there fast.
  • You can evaluate systems with fuzzy success criteria and design experiments that give reproducible answers.
  • You treat data as infrastructure. You've owned an annotation pipeline, an eval set, or a feedback loop, and you know label quality drives model quality more than architecture does.
  • Excellent ethics and judgment on authorization, customer safety, responsible disclosure, and offensive capability. You'll build technology that is dangerous if misused. We need someone who takes that seriously without freezing, and who will push back when a customer asks for something we shouldn't do.
  • You work without much direction. You'll pick the problems as often as you're handed them.

Nice to have

  • Security background of any kind: offensive or defensive work, CTFs, homelabs, research, or time at a security company. We strongly prefer this and will teach the rest.
  • Social engineering, phishing infrastructure, or human-layer security
  • Deepfake detection or media forensics, which is this problem from the other side
  • Trust and safety, model safety, or abuse prevention, especially where controls had to survive a motivated attacker
  • Content provenance and watermarking, C2PA or otherwise
  • Having built under real regulatory constraint, where compliance shaped the architecture
  • Telephony and SIP, or voice agents in production
  • Fine-tuning, distillation, or self-hosted inference at scale
  • Published research, open source, CVEs, or conference talks
  • Multi-tenant SaaS, and the isolation problems that come with running adversarial work for many customers at once

What success looks like in 90 days

  • A real-time voice simulation running end to end in a live customer engagement, not a demo
  • A measurable gain in realism or success rate, with the evaluation method to prove it
  • A labeling and evaluation pipeline our researchers actually use, so finished engagements become structured data instead of call recordings and memory
  • Safety and authorization architecture documented, enforced in code, and attacked by our own researchers
  • A clear view on the applied AI roadmap for the next two quarters, including what we build and what we buy

Why this role is unusual

Almost nobody gets to work on this. Offensive AI research usually happens either in academia with no real targets, or inside organizations that will never publish. Here you build it, run it against real companies that asked you to, watch what happens, and publish what you learn. The feedback loop is the point.

See how well your resume matches this job before you apply

Run a free ATS check