The archive / 2026

Video talk Hosted by Stan

LLMs Don't Hack, They Guess

Friday, March 20, 2026 · Stan, Toronto ·43 attended

Speakers

Talk night at Stan, who also sponsored. Jeet, who does offensive security at Robinhood, spent a year pointing LLMs at real production code and open source. He said simple patterns are easy for AI, with Claude Code finding a pickle.loads hidden in a 10,000-line codebase in two or three minutes.

The hard case was a real Helm CVE. The bug spans about 20 layers of code, and Cursor and Claude both turned up false positives and missed the actual bug, about 65 cents all wasted. A scoped multi-agent pipeline of a source-to-sink tracer, a threat modeler and a vulnerability finder got further than one broad agent, finding the file write sink and the right abuse case, though it still missed the final exploit step.

#ai#offense#code-review

Sources & references