Input Sanitization for Agentic Systems: What Actually Works
Speakers
-
KC Udonsi
Head of Cybersecurity at Stan. DC416 co-organizer. Building with agentic systems since 2022.
KC Udonsi, Stan’s head of cybersecurity and a DC416 co-organizer, covered the semantic firewall he built after the AI policy he wrote at Stan got clicked through and ignored. It sat as a proxy in front of the model, scanning input for PII and prompt injection and checking responses on the way out.
To protect names, it swapped Jane Cooper for Alex Morgan and mapped the real name back afterward. Placeholders like [NAME_1] had led the model to invent names. For injection, four detectors ran in parallel (regex, a local DeBERTa classifier, semantic drift detection and an LLM judge) and a majority vote blocked the input. Stan sponsored the night at its Toronto office.
Sources & references