Silence of the RAM: constant-memory safety probing for long-context LLMs. A streaming hard-max probe with a measured Theta(min(C,N)) memory law (flat 19.1 MiB to N=131,072 vs 652.1 MiB for softmax pooling), the subgradient obstruction it creates, and the fragmentation attack that defeats it.
machine-learning pytorch ai-safety calibration-methods mechanistic-interpretability long-context-llm activation-probing cascading-classifiers
-
Updated
Sep 10, 2026 - Python