Foundation Models / Cybench Evaluation
Small security models reach the Cybench frontier.
In Singularity's 25-model Cybench evaluation, the 1B and 3B models challenged systems hundreds of times larger. GPT-5.6 Sol leads the comparison, while GLM-5.3 raises the open-model ceiling. The consequential result is capable security reasoning that can stay inside the network it is meant to defend.
The short answer
What do the Cybench results show?
GPT-5.6 Sol with xhigh reasoning scored 84.0, while the open-weight GLM-5.3 scored 78.0. Singularity-3B reached 72.0 with 248 times fewer parameters than GLM-5.3, and Singularity-1B scored 64.0. Within this evaluation, compact models did not merely trade size for performance. They redrew the efficiency frontier.
Figure 1 / Interactive
Model size tells only part of the story.
Move up for a higher score and left for a smaller model. Use the controls to isolate vendors, labels and the measured frontier.
Singularity-3B scored 72.0, six points behind GLM-5.3 at 78.0 and 744B parameters. The ratio describes parameter count, not measured latency or total serving cost.
GPT-5.6 Sol with xhigh reasoning leads the 25-model comparison. OpenAI does not disclose its parameter count, so it appears in the score-only field.
Singularity-1B outscored every open comparison model except GLM-5.3, including GLM-5.2 at 744B parameters.
Why it matters
Why small security models matter
Security teams work with highly sensitive data. A compact model can bring capable reasoning inside the network, close to the evidence, without requiring that telemetry travel to an outside cloud.
Prompts, telemetry and model weights can remain on-premises or in a fully air-gapped environment.
The model can work beside sensors, tools and investigation data instead of receiving a distant copy.
Fewer parameters can make private deployment more practical without giving up useful benchmark performance.
What the result means: model size still matters, but it did not determine performance by itself in this evaluation.
Results for all 25 models
| Model ⇅ | Vendor ⇅ | Params (B) ⇅ | File F1 (/100) ⇅ | Access ⇅ |
|---|
How to read the result
- Higher File F1 is better. Smaller parameter counts appear farther left.
- The mint field marks the compact Singularity models in this comparison.
- The dashed curve connects Singularity-350M, Singularity-1B and Singularity-3B.
- Closed models appear in a separate score-only field because their parameter counts are not public.
Important scope: File F1 is Singularity's internal reporting metric, not the official Cybench leaderboard metric. These results use one internal harness on the public task split and should be interpreted within that setup.
Methodology
Raw File F1 is rescaled to 0 through 100, with 0.30 raw F1 represented as 100. The same harness and scoring method were applied to all models shown. Open-model parameter counts come from published model cards. Closed models are compared by score only, with parameter count intentionally omitted.
The public Cybench framework contains 40 professional capture-the-flag tasks from four competitions across cryptography, web security, reverse engineering, forensics, miscellaneous challenges and exploitation.
Evaluation date: August 2026
Published: September 4, 2026
Models shown: 25, including one preview model
Primary sources
Direct answers
Questions about the evaluation
What is Cybench?
Cybench is a public framework for evaluating the cybersecurity capabilities and risks of language-model agents. It includes 40 professional capture-the-flag tasks from four competitions, organized across six security categories.
How did Singularity models perform in this evaluation?
Singularity-1B scored 64.0 and Singularity-3B scored 72.0 on the normalized File F1 scale. GLM-5.3 scored 78.0 at 744B parameters, while GPT-5.6 Sol with xhigh reasoning scored 84.0.
Is File F1 the official Cybench leaderboard metric?
No. File F1 is Singularity's internal reporting metric for this evaluation. One harness and one normalization method were applied across the models shown, and the results should be read within that scope.
Why compare parameter count?
Parameter count is an imperfect but useful indicator of model scale. It does not directly measure latency, memory use or serving cost, so this page presents it as context rather than a complete efficiency measure.
Can Singularity Foundation Models run in an air-gapped environment?
Yes. The models in this evaluation are designed for on-premises and fully air-gapped deployment, allowing prompts, telemetry and model weights to remain inside the operator's environment.
Bring security reasoning to the data, without sending the data away.