The Sthenos Cyber Division is the security testing practice of Sthenos Technologies. It attacks a client’s systems the way a real adversary would, through penetration testing, red teaming and AI and LLM red teaming, and then measures whether the client’s defenses would have seen it, through adversary emulation and detection validation. Every engagement starts with signed written authorization, is mapped to NIST SP 800-115, the OWASP testing guides and MITRE ATT&CK, and ends in an evidence-backed report, with a retest available to confirm each fix.

ServicesNine

From a fast attack-surface check to a full purple-team program.

MethodologyNIST 800-115

With OWASP WSTG, PTES and MITRE ATT&CK for coverage and reporting.

AuthorizationSigned first

No scan, test or simulation starts before written authorization is in hand.

ApproachOffense and defense

We break in, then show which of your alerts fired and which stayed silent.

Built quietly, now taking engagements

The division built and verified its lab before taking a single client: recon, web and cloud testing, AI red teaming, adversary emulation, detection engineering and consented mobile forensics, each stood up and proven against deliberately vulnerable targets we own. It is now open for engagements with small and mid-sized businesses, enterprise security teams and public-sector buyers.

Talk to the Cyber Division

What the Cyber Division does

Most security vendors sell one half of the picture. A penetration test tells you whether an attacker can get in. A monitoring product tells you what your tools logged. Neither tells you the thing a board actually wants to know: when a real attacker does these specific things, would we notice? The Cyber Division runs both halves as one loop, so every engagement ends with a measured answer rather than a list.

One loop, offense into defense

The Sthenos Cyber Division engagement loop Six stages in sequence: discover the attack surface, attack it under authorization, check what your defenses detected, measure coverage against MITRE ATT&CK, report with evidence, and retest each fix, which feeds back into discovery. Discoverattack surface Attackunder authorization Detectdid an alert fire? MeasureATT&CK coverage Reportevidence, fixes Retestconfirm each fix The loop repeats on a schedule for clients on continuous testing OFFENSE DEFENSE ASSURANCE

Discovery feeds every offensive service. Whatever the attack produced is checked against your detection stack, recorded against MITRE ATT&CK, written up with evidence, and retested after you fix it.

Our nine security services

Each service below names what you walk away with and whether the work is run by tooling with a human reviewing it, or performed by hand by a human tester. A scan is never sold as a penetration test, and the report says plainly which one you bought.

Attack-surface check and vulnerability scan

What an attacker can see from the internet, and which of it is obviously exposed. An asset inventory and a severity-ranked findings list.

Automated, human-reviewed

Web application penetration test

A tester attacks your application by hand for injection, broken access control and business-logic flaws that scanners miss.

Human-led

External network penetration test

Proves what someone on the public internet can actually do to your perimeter, with controlled exploitation and evidence.

Human-led

Internal and Active Directory test

From a foothold inside, how far an attacker gets and how fast they reach domain admin, told as an attack-path narrative.

Human-led

Cloud posture and configuration review

Microsoft 365 and Entra, AWS or Google Cloud scored against CIS Benchmarks, with a prioritized fix roadmap.

Automated, analyst-written

Phishing and social-engineering simulation

Measures whether your people would stop a real attacker, with your HR and legal sign-off, and turns the result into training.

Automated delivery, human design

AI and LLM red teaming

Prompt injection, data leakage, jailbreaks and unsafe agent actions in your chatbot, agent or RAG system, mapped to the OWASP LLM Top 10.

Human-led

Continuous monitoring and retesting

Recurring scans that report what is new, what changed and what regressed, plus proof that old findings stayed fixed.

Automated, human triage

Adversary emulation and detection validation

Real attacker techniques run against your estate, technique by technique, to show which alerts fired and which stayed silent.

Human-led purple team

Penetration testing

A penetration test puts a qualified attacker against your systems on a fixed, signed scope and hands you confirmed, exploitable risk rather than a list of theoretical issues. The division tests web applications, APIs, external networks and internal networks including Active Directory, and follows NIST SP 800-115 for structure, the OWASP Web Security Testing Guide for web methodology and MITRE ATT&CK to describe every technique used. Tooling assists the tester with enumeration and scanning, but exploitation decisions are made by a person, and denial-of-service testing is never run unless it is separately and explicitly authorized.

Scoping, testing types, black box versus grey box, pricing drivers and what each compliance framework asks for are covered in full on our penetration testing services page.

AI and LLM red teaming

AI red teaming is adversarial testing of an AI or large-language-model application: a chatbot, an agent that takes actions, a retrieval-augmented (RAG) system or an AI feature inside your product. The goal is to find the ways it can be manipulated into leaking data, ignoring its instructions, producing harmful output or taking actions it should not, before a customer or an attacker does. It combines classic application testing with attack techniques that only exist for AI systems.

What we test for

  • Direct and indirect prompt injection, including instructions hidden in documents your system retrieves
  • Sensitive-information disclosure and system-prompt extraction
  • Jailbreaks and off-policy output
  • Insecure tool, plugin and agent-action use, and excessive agency
  • Supply-chain and RAG-poisoning exposure

How results are expressed

OWASP Top 10 for LLM ApplicationsThe risk taxonomy every finding is mapped to.
MITRE ATLASAdversarial tactics and techniques against AI systems.
NIST AI RMFThe governance frame for the remediation advice.

Breadth comes from two open-source tools stood up and verified in our lab: garak, NVIDIA’s LLM vulnerability scanner, and PyRIT, Microsoft’s AI red-team framework. Depth comes from the tester, who chains attacks and judges real impact. Findings map to the OWASP Top 10 for LLM and GenAI and, where relevant, MITRE ATLAS, with governance framing from the NIST AI Risk Management Framework. Testing runs against a test instance of your system, never a production model you have not authorized.

Adversary emulation and detection validation

A penetration test asks “can we get in”. Adversary emulation asks “when an attacker does these named things, does your detection stack catch them”. We run real-world attacker techniques, one at a time or as a named-adversary chain, against your authorized estate and record for each one whether it was prevented, alerted on, logged without an alert, or missed entirely. It can run collaboratively with your defenders in the room, or as an unannounced exercise with a small trusted group on your side.

  1. EmulateRun an ATT&CK-mapped technique with controlled, non-destructive tooling.
  2. TelemetryConfirm what it left in your endpoint, network, identity and cloud logs.
  3. DetectCheck whether an alert fired, and how good it was.
  4. Tune and re-runWrite the missing detection, for example a Sigma rule, and re-run to watch it fire.

You get an ATT&CK coverage heatmap, a detection-gap report with the fix for each gap, tuned detection content validated by re-running the technique, and a prioritized hardening list. Engagements are sized in three tiers: focused validation of a set of techniques you care about, emulation of a named threat actor relevant to your sector using the MITRE Center for Threat-Informed Defense emulation library, or a recurring program that tracks your coverage improving over time.

Cloud, container and supply-chain security

A traditional network test walks straight past a misconfigured cloud account, an over-permissioned identity, a vulnerable container image, a leaked key in source control or a vulnerable dependency deep in the build. The division covers each of them.

Cloud posture

Microsoft 365 and Entra, AWS or Google Cloud, scored against CIS Benchmarks through time-limited, read-only access. Prowler and ScoutSuite collect, an analyst interprets.

Kubernetes

Cluster benchmarking with kube-bench and exposure checks with kube-hunter.

Containers and SBOM

Image and dependency scanning with Trivy and Grype, and a software bill of materials with Syft.

Code and secrets

Static analysis with Semgrep and secret scanning with gitleaks and TruffleHog.

A cloud posture review uses read access, not attack, so it generally does not need your provider’s testing authorization; if you add an attack component it becomes a penetration test and follows those rules instead. If you build software with us, the same scanners can run inside your own CI pipeline as pre-merge gates.

Phishing and social-engineering simulation

A controlled phishing campaign against your own staff, run only with written authorization from an accountable executive and sign-off from your HR and legal teams, measures who clicks, who submits credentials and, just as important, who reports it. Email phishing is the default; voice, SMS or physical pretexting happen only if they are explicitly scoped. Results are for improvement, not blame, and are reported in aggregate unless your own policy and consent framework say otherwise.

Continuous monitoring and retesting

A once-a-year test leaves eleven months in which a new subdomain, a changed firewall rule or a regressed fix goes unnoticed. Continuous testing re-runs attack-surface and vulnerability checks on your agreed assets on a schedule, reports what is new, what changed, what was fixed and what came back, and confirms previously reported findings are still closed. Anything that needs a human is escalated to a tester rather than left in a dashboard.

Consented mobile spyware checks

For people who believe their phone may be targeted, the division checks a device owner’s own backup for indicators of mercenary spyware of the Pegasus and Predator class. It uses the Mobile Verification Toolkit, built by Amnesty International’s Security Lab, against public indicator feeds. This is defensive forensics done only with the owner’s consent and on devices they own or the client owns. Sthenos detects this kind of spyware; it never deploys it.

The lab and toolchain behind the work

Every service above runs on tooling the division has installed, version-pinned and proven in its own lab against deliberately vulnerable targets it owns, such as OWASP Juice Shop and DVWA, before it is ever pointed at a client. The stack is drawn from the current open-source frontier rather than one vendor’s suite, so each tool can be swapped when something better appears.

cyber-division / toolchain
recon and attack surfacesubfinder, dnsx, httpx, naabu, katana, Amass, nuclei
ai and llm securitygarak, PyRIT
adversary emulationMITRE Caldera, Atomic Red Team, CTID Adversary Emulation Library
detection engineeringWazuh, Suricata, Sigma
purple-team trackingVECTR, ATT&CK Navigator
cloud and kubernetesProwler, ScoutSuite, kube-bench, kube-hunter
supply chain and codeTrivy, Syft, Grype, Semgrep, gitleaks, TruffleHog
consented phishingGoPhish, with a mail sink that cannot reach the internet in the lab
mobile forensicsMobile Verification Toolkit with public STIX2 indicator feeds

Alongside that toolchain, the division is developing AI-assisted testing agents, still being benchmarked, that run on models hosted on hardware Sthenos owns. Where they are used, they propose and sequence steps, a human approves each exploitation step before it runs, and the AI side of the work does not send your target data to a third-party AI service.

How an engagement runs

  1. Scope and authorizationAssets, IP ranges, testing windows, rules of engagement, emergency contacts and a stop procedure are agreed and signed. Nothing touches your systems before this.NIST SP 800-115 planning
  2. ReconnaissanceWe map what is reachable: domains, subdomains, live hosts and exposed services, inside the signed scope only.Discovery
  3. Scanning and enumerationPorts, services, versions, content and configurations are enumerated and merged into one target map.Discovery
  4. Vulnerability analysisCandidate weaknesses are assembled from scanners and manual review, and a human clears the false positives.OWASP WSTG
  5. ExploitationControlled exploitation to the depth you authorized, each step approved by a person before it runs.Attack
  6. Post-exploitationWhere in scope, privilege escalation and lateral movement to show real business impact, recorded step by step.MITRE ATT&CK
  7. ReportingFindings with severity, reproduction steps, evidence and remediation, plus an executive summary and a debrief with your engineers.Reporting
  8. RetestThe same tool that proved each finding is re-run against your fix, and every finding is marked fixed, partially fixed or still open.Verification

What lands on your desk

  • A findings report: severity, reproduction steps, evidence and remediation for every issue
  • An executive summary written for non-technical leadership
  • An attack-path narrative for internal and Active Directory tests
  • An ATT&CK coverage heatmap for emulation and purple-team work
  • Tuned detection rules, validated by re-running the technique
  • A CIS-benchmarked posture score and fix roadmap for cloud reviews
  • A technical debrief to hand findings to your team
  • A retest result for each finding: fixed, partially fixed or open

The rules we work by

Written authorization first

No scanning, testing, phishing or exploitation begins until an accountable owner of the target has signed off on scope, ranges, windows and rules of engagement.

Inside the scope, always

Anything outside the signed scope is not touched. For cloud tenants and hosted platforms you do not own outright, the provider’s authorization is obtained where it is required.

No inflated credentials

We do not claim blanket certifications. The proposal names the tester assigned to your engagement and the credentials your requirement calls for.

Consent for people and devices

Phishing needs your HR and legal sign-off. Mobile checks need the device owner’s consent. We detect spyware; we never deploy it.

Standards every result is measured against

If you are building software that has to pass an audit, our guides to FedRAMP-ready development and PCI DSS-ready development cover the build side. Government buyers can find our registrations on the capability statement.

Frequently asked questions

What is the Sthenos Cyber Division?

It is the security testing practice of Sthenos Technologies, a US software and AI company in Tysons, Virginia. The division runs offensive testing, including penetration tests, red teaming and AI and LLM red teaming, together with defensive validation, including adversary emulation and detection engineering, so a client learns both whether an attacker can get in and whether their defenses would notice.

What is the difference between a vulnerability scan and a penetration test?

A vulnerability scan uses automated tools to find known weaknesses and exposed services quickly and broadly, without exploiting anything. A penetration test is performed by a human tester who confirms which weaknesses are actually exploitable, chains them together and finds logic flaws that scanners cannot see. We sell both and every report states plainly which one was performed.

What is AI red teaming?

AI red teaming is adversarial testing of an AI or LLM application, such as a chatbot, agent or RAG system, to find prompt injection, data leakage, jailbreaks and unsafe actions before attackers do. We map every finding to the OWASP Top 10 for LLM and GenAI and, where relevant, to MITRE ATLAS, and test against a test instance of your system.

What is adversary emulation, and how is it different from a penetration test?

A penetration test asks whether an attacker can get in. Adversary emulation runs specific real-world attacker techniques, mapped to MITRE ATT&CK, and measures technique by technique whether your detection stack prevented, alerted on, logged or missed each one. The result is a coverage heatmap and the detection rules to close each gap.

Do you need written authorization before testing?

Yes, always. No scanning, testing, phishing or exploitation begins until an accountable owner of the target signs written authorization naming the scope, the IP ranges or domains, the testing window and the rules of engagement. Phishing additionally needs your HR and legal sign-off.

Do you use AI in your testing, and where does our data go?

Our core toolchain is conventional security tooling run by people. The division is also developing open-source AI-assisted testing agents, still being benchmarked; where they are used, a human tester approves each exploitation step before it runs, and because the models run on hardware Sthenos owns, the AI side of the work does not send your target data to a third-party AI service.

Can you check whether my phone has Pegasus spyware?

Yes, with your consent and on a device you own. We analyze your own device backup with the Mobile Verification Toolkit, built by Amnesty International’s Security Lab, against published indicators for mercenary spyware of the Pegasus and Predator class, and report what was and was not found.

Are your testers certified?

We do not make blanket certification claims. Each proposal names the tester assigned to the engagement and the credentials your requirement calls for, such as OSCP or GPEN, so you can verify the person doing the work rather than a logo on a website.

How much does a security engagement cost?

Every engagement is quoted on its scope: the number of applications, hosts, cloud accounts, AI interfaces or employees in scope, and the depth of testing you authorize. Tell us what you need tested and what is driving the requirement, and we will come back with a written proposal.

Ready to test what an attacker would see? Tell us what is in scope and what is driving the requirement, and the Cyber Division will come back with a proposal that names the tester, the standards and the retest terms. Talk to the Cyber Division or call +1 301-793-3980.

Contact us
Talk to an engineer

We’re happy to answer any questions you may have and help you determine which of our services best fit your needs.

Rates and delivery
What happens next?
1

We schedule a call at your convenience

2

We run a short, bounded discovery, scoped per engagement

3

We give you a costed roadmap before committing to a build

Request a Free Consultation
Book a 30-minute call →Prefer to talk first? Skip the form and grab a time directly.

By submitting this form you agree to our privacy policy.

We aim to reply within one business day