Apps & Tools
AgentRE-Bench
A deterministic benchmark for AI agents reverse-engineering Windows PE binaries
PlayableOpen source
AgentRE-Bench media is blocked
Allow external media to connect to the provider and play this content.
Description
A research framework evaluating how frontier AI agents perform static analysis on paired stripped and unstripped Windows binaries. The benchmark uses deterministic scoring across 120 evaluation episodes to track model performance in identifying behavior, recovering configurations, and tool planning without relying on LLM judges.
Descriptions, tags, and model credits may be AI-generated or inferred from public sources and can be incomplete or wrong. Learn more.