
Open-source endpoint security CLI for monitoring, detecting, investigating, and optionally blocking risky AI coding-agent activity.
Do not bounce yet
Read the fit check, compare one alternative, then decide whether the vendor page is still your best next click.

Quick Verdict
Make the fit call first. Vendor pages are good at selling, but they rarely tell you where the product is a bad match.
Compare Next
This is where visitors usually jump out too early. Read one deeper take or open one alternative so the next click is informed instead of impulsive.
Alternative profile
MIT-licensed local memory server and CLI for carrying coding context, summaries, and handoffs across Claude Code, Codex, Cursor, OpenCode, and other agent clients.
Alternative profile
MIT-licensed codebase knowledge graph and MCP server for structural search, natural-language exploration, data-flow tracing, and agent-assisted edits across mixed-language repositories.
Alternative profile
MIT-licensed control plane for running Claude Code, Codex, Cursor, Grok, and OpenCode threads from desktop, web, iOS, and Android clients.
Numbat is an open-source endpoint security layer for teams adopting coding agents faster than their existing security tooling can observe them. Instead of wrapping one model or IDE, it connects to supported agent hooks, telemetry, and stored session artifacts, normalizes the activity, and evaluates a shared set of detection rules locally. That makes it relevant to serious vibe coding environments where Claude Code, Codex, Cursor, OpenClaw, or other harnesses can run commands and touch sensitive developer systems.
Numbat is Perplexity's local-first security and forensics layer for coding agents. Its static Go binary discovers supported agent installations, normalizes live hooks, OTLP logs, and on-disk session artifacts into one event model, evaluates built-in or custom CEL rules, and can reconstruct timelines or build verifiable case bundles. Synchronous pre-action blocking is opt-in, limited to supported hooks, and disabled in every shipped rule by default. The important caveat is that Numbat is an early v0.1.x project: coverage varies by agent and surface, several current SQLite session stores are deferred, findings are rule matches rather than proof of compromise, and records can retain sensitive endpoint context even after redaction.
Choose Numbat when you need one local event model across several coding-agent harnesses rather than separate ad hoc audit scripts for each tool.
Its read-only inventory and forensic scanning commands provide value before you install hooks or enable any blocking behavior.
The opt-in enforcement path is useful for carefully tested high-risk actions, but shipped rules remain monitor-only and host coverage must be checked first.
Do not treat it as a finished universal sandbox: the project is early, durable-store coverage has gaps, and sensitive output still requires strong access controls.
Discovers supported local coding-agent installations and scans their on-disk session artifacts without executing commands found in those records.
Normalizes lifecycle hooks, generated plugins, OTLP/HTTP logs, and stored artifacts into versioned NDJSON events and findings.
Evaluates built-in and custom CEL rules, including multi-step sequences such as secret access followed by outbound transfer.
Supports opt-in pre-action denial on compatible agent hooks while keeping all shipped rules monitor-only by default.
Reconstructs per-session timelines and produces portable case bundles with SHA-256 manifests for investigation workflows.
Ships one static Go binary for macOS, Linux, and Windows on amd64 and arm64.
Use read-only discovery to see which supported agents and local data surfaces exist before changing hook configuration.
Normalize supported hooks, plugins, and OTLP logs into a common event stream for local CEL-based detection.
Scan stored session artifacts, build a timeline, and export a manifest-backed case bundle without executing recorded commands.
Promote selected operator rules to enforcement only after monitor-mode validation and verification that the target agent supports a synchronous deny path.
Security and platform teams rolling out coding agents on developer endpoints
Developers who need inspectable audit trails across multiple agent harnesses
Incident responders reconstructing risky or unexpected agent actions
Open-source evaluators comparing agent guardrails, telemetry, and endpoint controls
Inventory coding agents on developer laptops and identify which live and forensic surfaces can be monitored.
Monitor Claude Code, Codex, Cursor, OpenClaw, Hermes, and other supported harnesses with a shared event and rule model.
Investigate an agent incident by reconstructing prior sessions from local artifacts and exporting a verifiable case bundle.
Pilot narrowly scoped pre-action rules for dangerous commands after validating monitor-only findings and each host's enforcement semantics.
Invariant Guardrails
Semgrep
Socket
Endor Labs
MIT-licensed local memory server and CLI for carrying coding context, summaries, and handoffs across Claude Code, Codex, Cursor, OpenCode, and other agent clients.
MIT-licensed codebase knowledge graph and MCP server for structural search, natural-language exploration, data-flow tracing, and agent-assisted edits across mixed-language repositories.
MIT-licensed control plane for running Claude Code, Codex, Cursor, Grok, and OpenCode threads from desktop, web, iOS, and Android clients.
Strong picks usually survive one more internal check. Read deeper, compare a neighbor, then leave for the vendor page if the fit still holds.