Can an Open Model Do Security Research? Cantina’s apex-flash-1 Solves 40 of 60 Held-Out Bug Tasks

Cantina Security, working with Yeta Labs, has released apex-flash-1. This is an open-weights model trained specifically for vulnerability research. It is a…

By Vane October 5, 2026 2 min read

Cantina Security, working with Yeta Labs, has released apex-flash-1. This is an open-weights model trained specifically for vulnerability research. It is a reinforcement learning fine-tune of Z.ai’s GLM-5.3-Flash, released on Hugging Face under the MIT license.

The weights are deployable on vLLM, SGLang, or Transformers, but BF16 requires roughly 640 GB of GPU memory.

What Cantina Built

apex-flash-1 has 321.3B total parameters, per its Hugging Face safetensors metadata. The GLM-5.3-Flash base is a Mixture-of-Experts model with 18B active parameters.

Cantina trained it with GRPO using a rank-256 LoRA plus selective full-parameter training. The data covers 150 tasks built from 50 real vulnerability cases.

Each case appears in 3 variants: guided whitebox, focused whitebox and focused blackbox.

Authorization, identity and scope flaws make up 72% of cases. Accounting and numerical precision bugs add 18%. Time validation, business rules and SSRF cover the rest.

As per the model card on HF, RL rollouts ran inside the Codex agent harness on production-like software and protocol environments.

Benchmark Results

Cantina evaluated 60 tasks from 20 held-out vulnerability cases. Each model ran the set once, with costs estimated from provider pricing.

  • apex-flash-1: 40/60 solved (66.7% pass@1), about $2.38
  • GLM-5.3-Flash (base): 36/60 solved (60.0%), about $4.56
  • Claude Opus 5 High: 43/60 solved (71.7%), about $74.68

Opus solved 3 more tasks but cost about 31x more per run. That is roughly $0.06 per solved task for apex-flash-1 versus $1.74 for Opus. These are company-reported numbers on an internal benchmark.

A Worker Model, Not an Orchestrator

Cantina positions apex-flash-1 as a worker orchestrated by a larger model. The card lists code reading, tool use, exploit development and verification as target skills.

An experimental apex-flash-1-abliterated variant ships with modified refusal behavior. It was not separately evaluated.

Cantina’s rationale is that defenders need capable models they can run and control locally.

How It Compares

Featureapex-flash-1Aikido Altar-1Cisco Foundation-Sec-8B-ReasoningGLM-5.3-Flash
DeveloperCantina Security + Yeta LabsAikido SecurityCisco Foundation AIZ.ai
Base modelGLM-5.3-FlashGLM-5.3 (pruned)Llama 3.1 8BOwn pretraining
Size321.3B total, BF16328 GB, INT4 (W4A16)8B320B total, 18B active
LicenseMITInherits GLM-5.3 licenseCustom (see NOTICE.md)MIT
Security methodGRPO RL on 50 real vulnerability casesExpert pruning (REAP) + quantizationInstruction tuning + RLHF on security QAGeneral-purpose base
Primary useAgentic vuln research workerAir-gapped autonomous pentestingSOC triage and threat defenseGeneral coding and agents
HardwareMulti-GPU node (~642 GB BF16 weights)4x H200 with vLLMSingle GPUMulti-GPU node
Published security result66.7% pass@1, 60 tasks60.4% recall, 32-CVE internal setCisco-reported security benchmarks60.0% on Cantina’s set

Sources: Cantina, Aikido, Cisco, Hugging Face model cards. © Marktechpost

What it means

For security teams, the immediate value is cost reduction. Running a full benchmark suite with apex-flash-1 costs about $2.38 compared to roughly $74.68 for a comparable high-end API model. The model allows organisations to run vulnerability research in-house on local hardware rather than relying on external cloud services.

Scroll to Top