Aikido Security Releases Altar-1: An Open-Weight Security Model Pruned From GLM-5.3 to 328 GB

Aikido Security has released Altar-1, its first open-weight security model derived from Z.AI’s GLM-5.3 and compressed to 328 GB. The weights are…

By Vane September 25, 2026 3 min read
Aikido Security Releases Altar-1: An Open-Weight Security Model Pruned From GLM-5.3 to 328 GB

Aikido Security has released Altar-1, its first open-weight security model derived from Z.AI’s GLM-5.3 and compressed to 328 GB. The weights are available on Hugging Face and run on vLLM using a single node of four NVIDIA H200 GPUs. The model powers Aikido Machine, an autonomous pentesting appliance designed for on-premises and air-gapped networks.

The Problem: Security Context Cannot Leave the Network

Closed frontier models operate on infrastructure owned by third parties. Deploying them risks sending source code, architecture documents, and unremediated security findings outside the network perimeter. This restriction affects banks subject to data-residency mandates and OT operators with no internet connection.

Open-weight models address the residency issue but introduce a deployment gap. Mixture-of-experts (MoE) architectures require storage for every expert, even when a specific workload activates only a fraction of them. Security agents also maintain long-running context, meaning the key-value cache competes directly with model weights for the same GPU memory.

How Altar-1 Was Built

GLM-5.3 is a 753B parameter MoE model where each token routes to eight of the 256 available experts per layer, resulting in approximately 40B active parameters. Aikido applied two compression steps to reduce this footprint:

  • Step 1: Quantization: The process started from the cyankiwi GLM-5.3-AWQ-INT4 checkpoint. AWQ stores routed expert weights in 4 bits with 16-bit activations (W4A16). Attention, the shared expert, dense layers, and the head remained in BF16.
  • Step 2: Expert pruning: Aikido used Cerebras REAP (Router-weighted Expert Activation Pruning). This method scores each expert based on router weight and output magnitude, rather than just selection frequency. Altar-1 retains 168 of the 256 routed experts per layer and removes 88 (34.4%). No retraining was involved.

Calibration relied on traces from Aikido’s pentesting harness, alongside coding, tool calling, reasoning, and multilingual Wikipedia text. The company states no customer data was used. Each expert was scored by its largest share of any single domain’s routed work. This approach protects specialist experts handling code, rare languages, and structured output.

Routing logic remains unchanged. The router still selects eight experts per token, now drawn from a pool of 168, maintaining about 40B active parameters.

CheckpointStored weights
GLM-5.3, BF161,506.7 GB
GLM-5.3, AWQ INT4488.2 GB
Altar-1, pruned W4A16328.0 GB

Altar-1 is 78.2% smaller than the BF16 version and 32.8% smaller than the AWQ parent. On fidelity, the model shows a KL divergence of 0.506 nats against full BF16 on a sealed 25-prompt panel. An EXL3 build of the same cut scores 0.511. Further details appear in the public fidelity study.

Benchmark Results

The Aikido team tested Altar-1 on an internal CVE benchmark covering 32 known vulnerabilities across 30 repositories, with three runs per case.

ModelAvg recall per runFound at least once
GLM-5.3, BF1665.6%25 of 32
GLM-5.3, AWQ INT461.5%23 of 32
Altar-160.4%23 of 32

Compared with the AWQ checkpoint, pruning cost about one point of recall with no change in coverage. Against the parent model, Altar-1 maintains coverage of 23 of 25 identified vulnerabilities (92%) while recording 5.2 points lower recall.

The benchmark scope is narrow. It measures targeted CVE rediscovery inside a pipeline that uses other models for surrounding stages. It does not measure blind discovery, exploit validation, or fix proposals. Aikido also reports that Altar-1 found a valid critical-severity vulnerability during a client’s production pentest. This is a single result reported by the vendor.

Deployment and License

The model card requires Hopper GPUs (H100 or H200). Aikido notes that 328 GB across four H200 GPUs leaves room for a 128k-context KV cache at production batch sizes.

vLLM selects the Marlin MoE kernel automatically. A four H100 80 GB node has only 320 GB of memory, which is less than the 328 GB required for the weights.

Altar-1 inherits the GLM-5.3 License. The license permits commercial use, modification, and redistribution. Model-as-a-Service operators with more than $10B in revenue over 12 months must first pass a Z.AI security review. Altar-1 is open-weight, not OSI-approved open source.

Altar-1 also powers Aikido Attack, AI Code Analysis, and Deep Review. Next, Aikido plans to try lower-bit formats like EXL3 so it can keep more experts. It also plans to fine-tune models for security workflows.

What it means

Organisations with strict data sovereignty rules can now deploy a capable security model on their own hardware without sending data to external clouds. The trade-off is a reduction in recall from 65.6% to 60.4%, but the model retains 92% of the original coverage. Hardware requirements are specific: four H200 GPUs are needed to hold the weights and context simultaneously. Operators generating over $10B in annual revenue face a review hurdle before commercial deployment.

Scroll to Top