← Back to all fieldnotes
Detection engineering

Reviewing Falco Noise with Local AI

The problem

The trap with Falco noise is that the alerts are often your own infrastructure: health checks, cron jobs, backup scripts, deployment commands. The rule correctly matches the event. The engineer still has to decide whether it needs attention.

A broad exception can clear the queue and hide the next real problem at the same time. I wanted a review loop: give the tool existing rules, alert logs, and workload authorization context; see what it would change; accept or reject each proposal.

That became auto-falco-rule-builder. Its main use is tuning existing detections from evidence. Repetition alone is never a reason to suppress an alert.

The review loop

The CLI groups alerts by their observed fields. A local model assesses each group as keep alerting, likely noise, or uncertain, with a reason and event count. Those labels are judgments to review, not verified incident or authorization decisions.

It proposes an exception only for groups assessed as noise. Before approval, I see the change in plain English, why the model recommends it, the exact field values, the number of recorded alerts affected, the future detection tradeoff, and the YAML diff.

Accept, reject, feedback, or quit. One rule at a time, with no default approval. Feedback triggers reassessment; uncertain activity can trigger a question about the workload. Accepted changes are exported as replacement rule files, alongside diffs and session records. The originals stay unchanged. Nothing is installed by the CLI.

Where AI fits

Ollama runs qwen3.5:0.8b locally. Assessment and planning are separate calls. The model selects evidence groups and scope fields; Python checks the selection, derives exact values from the alerts, and compiles an exception macro into the existing rule condition.

The model has no shell, file-writing, or deployment tools. Calls use loopback-only Ollama with no cloud fallback. Selected command, account, and path values reach the local model; this is not blanket secret redaction. The older create and manual tune commands remain deterministic. AI generation of entirely new detections is not implemented.

The run

I built a lab with 500 fabricated alerts across ten rules: 400 events authorized by the scenario and 100 suspicious lookalikes. Each rule has changed commands, different accounts, and different parents that must remain detectable. A separate answer key scores the decisions; the model never reads it.

In my manual run, I accepted the health-check, artifact-fetch, and backup exceptions, then quit at the cache-scan proposal. The three accepted changes are projected to remove 120 noise events, leave 380 alerts, and retain all 100 suspicious events in the supplied records. That is a recorded-value comparison, not live alert reduction or runtime replay of those workloads.

The implementation has 83 passing tests. The manual run used pinned Falco engine validation and no replay. Earlier host-capture integration checks exercised detection, suppression, preservation, dependency order, and deliberate assertion failures. Those checks do not prove Kubernetes scope discrimination or suppression of the synthetic health-check activity.

The useful failure

The small model claimed an interactive-shell parent matched cron. Some assessments counted 42 likely-noise events while the proposed exception actually covered 40. The accepted exceptions still matched the intended scopes, but the explanations could have misled me into approving a different change.

Inspect the condition as well as the explanation. AI-assessed noise and events actually covered by an exception are different numbers. Valid YAML does not establish that an exception is justified.

A deeper review found a code limit too: the prompt asks for full command, parent, and account when available, but Python currently enforces only an activity field plus a context field. An executable-and-account exception could fit a small sample while hiding a different future command under that account. Enforcing every relevant authorization field in code is still hardening work to do. I would reject that broad scope.

Where it belongs

For an engineer onboarding Falco, the useful output is a reviewable change: what activity is expected, what the rule will stop reporting, and which evidence remains outside the exception. It gives a new engineer a practical route through logs and rule conditions.

For an organization, I would put it between alert triage and the detection-change pull request. Workload owners explain authorization; detection engineers inspect scope; the existing test and rollout process decides whether the change reaches production. Reports can contain private commands and account values, so they belong in that controlled workflow.

It helps review known activity. It cannot establish a clean-system baseline from alert frequency, replace an investigation, or guarantee preservation of future attacks. Future malicious activity matching an accepted exception will also be hidden by that rule.

Try it

From an installed checkout, with local Ollama and qwen3.5:0.8b available, start the synthetic lab:

python -u -m afb.cli review --rules-dir examples/ai-review-500/rules --logs examples/ai-review-500/alerts.synthetic.jsonl --context-file examples/ai-review-500/context.txt

This is a Docker-free static preview. Add --profile examples/ai-review-500/lab.profile.json --validation-mode container to require pinned engine validation. Compatible captures and explicit expectations enable real replay; the synthetic lab supplies neither. Failed validation or replay blocks acceptance. A static export is not runtime detection proof.

If you already load your own full rules through Helm, use the reviewed copies in accepted-rules/ as replacements in your existing rollout. Do not load them alongside the original full rules with duplicate names. Preserve the actual rule order and required dependencies.

Source and examples ↗ · Review guide ↗ · Run and score the lab ↗

The engineer's workflow

Supply rules, alert logs, and authorization context. Review the proposal and verification status before exporting it. The final rollout is a separate operator action.

Open workflow full size ↗

Inside the builder

Local AI recommends evidence and fields; Python checks and compiles the change. Source references are pinned to the implementation reviewed for this note. The diagram includes its current scope-check limitation.

Open component diagram full size ↗