Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,18 @@ The format follows [Keep a Changelog](https://keepachangelog.com/); versions fol

### Fixed

- **Patterns with a repeat inside a repeat are refused as too slow.** Patterns such as
`re:^(\w+)+$` used to pass the speed check, then took seconds on a single name or
address and could stall the network. They are now refused when you type them. A
pattern that cannot be built at all, such as `re:a{99999999999}`, is reported as an
invalid expression, and the command line exits with the configuration error code.

- **An expression that cannot be read no longer switches its field off during a
session.** When a destination, block or target expression could not be read as
settings were applied, the field was switched off, and a destination or target
switched off meant all traffic was impaired. The field now keeps its previous value,
and the log says so.

- **A command-line run stops on time even while its console window is paused.**
Selecting text in the console window pauses the program's output until the selection
ends. A run that reached its `--duration`, or hit a failure, used to wait for that
Expand Down
3 changes: 3 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -649,6 +649,9 @@ re:^ch.{1,8}e\.exe$ WRONG - it is split into "re:^ch.{1" and "8}e\.exe
* **`2000-1000` is an error** (reversed range), not an empty set.
* **A wildcard is not a regex.** In `chrome*` the star means "any run". In `re:chrome*` it means
"the letter `e` repeated 0+ times". If you write `re:`, you write a regex.
* **A repeat inside a repeat is refused as too slow**, for example `re:^(\w+)+$` or
`re:^([0-9:]+)+$`. Such a pattern can take seconds on a single name or address that almost
matches, and an address is checked on every packet. Write the repeat once: `re:^\w+$`.

Every syntax error is reported **immediately**: in the GUI the field turns red with the reason
beneath it (in the UI language), and the CLI ends with a readable `error: ...` - never a silent
Expand Down
124 changes: 84 additions & 40 deletions beantester/matchers.py
Original file line number Diff line number Diff line change
Expand Up @@ -42,6 +42,7 @@
the GUI can show it and the CLI can turn it into a clean error message.
"""
import fnmatch
import functools
import ipaddress
import re
import time
Expand Down Expand Up @@ -521,9 +522,13 @@ def _compare_predicate(op, number):
# The margin between "fine" and "not fine" is four orders of magnitude, which is
# why a clock is a fair judge here and why this does not flake on a slow runner.
#
# 🔴 The textbook example does NOT work: `^(a+)+$` is 0.001 ms at every length,
# because CPython optimises it away. A guard tested with it would prove nothing
# and look thorough. The two patterns above are the ones that actually blow up.
# 🔴 The textbook example is a trap in BOTH directions. `^(a+)+$` is 0.001 ms at
# every length on a run of `a` it MATCHES - the first way through succeeds - and
# this comment once concluded from that that it was harmless. It is not: it blows
# up when the run is followed by something the pattern refuses (refused at 20
# characters once `_REGEX_PROBE_TAILS` adds that, measured 2026-09-28). The ladder
# of matching runs let through every pattern of that shape, `^(\w+\s?)+$` among
# them, which spent 558 ms on ONE real process name.
# The budget covers the WHOLE trial, not each search in it, which is what keeps
# the cost of a refusal small: the ladder stops at the first length that has spent
# it, so a pattern is refused after roughly one growth step rather than after
Expand All @@ -542,74 +547,113 @@ def _compare_predicate(op, number):
# IPv4 (digits), process names (letters), IPv6 text (hex and colons, the shape the
# reproduction used), dotted names.
_REGEX_PROBE_UNITS = ("1", "a", "1:", "a.")
# Every run is tried as it is AND followed by a character none of the alphabets
# above contains. A repeat inside a repeat explodes when a long run it can consume
# is followed by something that makes the whole match fail, and a bare run always
# matches - so the ladder walked straight past `(a+)+$`, `^(\w+\s?)+$` and
# `^([\d:]+)+$`, the last one on the packet path at 2.7 s per packet on a real
# IPv6 address. MEASURED 2026-09-28 with the tail: all three refused by 20
# characters; the whole ladder for eleven ordinary patterns went from
# 0.013-0.025 ms to 0.021-0.044 ms; and a pattern near the budget that was refused
# 20 times in 40 is now refused 40 in 40.
# Known limit, said rather than hidden: a repeat that explodes only on ONE letter
# no unit contains (`(x+x+)+y`) still passes. That is not written by accident.
_REGEX_PROBE_TAILS = ("", "!")
# Up to 45, the longest IPv6 address in text form, which is the longest value an
# IP matcher is ever handed. A process name can be longer, and that is said out
# loud rather than covered badly: a pattern that is still fast at 45 characters
# and slow at 300 exists, and the capture-thread heartbeat in `engine.py` is what
# catches it.
# IP matcher is ever handed - and addresses and ports are the only values matched
# per packet. A process pattern runs on process NAMES (the resolver, the UI
# thread), and a name can be longer: the shapes that still pass the tail probe
# were measured at no more than ~1 ms per search at 256 characters (2026-09-28).
#
# Close steps at the bottom on purpose. The cost of a refusal is whatever the
# first over-budget length cost, so the rungs have to be near each other exactly
# where the explosion starts - `^((a*)*)*b$` is 7.6 ms at 8 characters and 1254 ms
# at 12, and a ladder that stepped straight from 8 to 12 would pay the second
# number to learn what the first already showed.
_REGEX_PROBE_LENGTHS = (6, 8, 10, 12, 14, 16, 20, 24, 32, 45)
# The ladder itself, built once. Lengths outer, alphabets inner, tails innermost:
# the ladder climbs for every alphabet at once, so a pattern that explodes on
# letters but not on digits is caught at the shortest length that shows it rather
# than after a full pass over the other.
_REGEX_PROBES = tuple((unit * length)[:length] + tail
for length in _REGEX_PROBE_LENGTHS
for unit in _REGEX_PROBE_UNITS
for tail in _REGEX_PROBE_TAILS)


def _blows_the_budget(rx):
"""True when climbing the ladder spends more than the budget.

Lengths outer, alphabets inner: the ladder climbs for every alphabet at once,
so a pattern that explodes on letters but not on digits is caught at the
shortest length that shows it rather than after a full pass over the other.
"""
"""True when climbing the ladder spends more than the budget."""
deadline = time.perf_counter() + REGEX_BUDGET_S
for length in _REGEX_PROBE_LENGTHS:
for unit in _REGEX_PROBE_UNITS:
rx.search((unit * length)[:length])
if time.perf_counter() > deadline:
return True
for probe in _REGEX_PROBES:
rx.search(probe)
if time.perf_counter() > deadline:
return True
return False


def _refuse_if_too_slow(rx, field, term):
"""Raise when a compiled pattern is too slow to sit on the packet path.
class _Refused(Exception):
"""A pattern that will not be used, and the ``errors.*`` key that says why."""

TWICE, and the second run is the one that decides, because a wall clock cannot
tell "this pattern burned five milliseconds" from "this thread lost the CPU for
five milliseconds". The obvious answer to that is a CPU clock, and it does not
work here: `time.get_clock_info("thread_time")` REPORTS a resolution of 1e-07
on this platform and MEASURES 15.625 ms (200 000 reads returned six distinct
values, 2026-09-02), which cannot see a 5 ms budget at all. A second run can:
a scheduling hiccup does not repeat in the same place, and backtracking does,
every time, deterministically. A good pattern never pays for this - it takes
the first run only, at 0.04 ms.
"""
if _blows_the_budget(rx) and _blows_the_budget(rx):
raise _err("errors.filter_regex_too_slow", field, term)
def __init__(self, key):
super().__init__(key)
self.key = key


def _compile_regex(pattern, field, term):
pattern = pattern.strip()
if not pattern:
raise _err("errors.bad_filter_regex", field, term)
@functools.lru_cache(maxsize=256)
def _accepted_regex(pattern):
"""The compiled pattern, once it has been judged fit to run per packet.

CACHED, and only what is accepted: ``lru_cache`` keeps nothing for a call that
raises. The judgement is a wall clock, so the same text could pass when the
form validated it and fail a moment later when "Apply" compiled it again - a
field that validated and then did not apply (external review, P3-13). Once
accepted, a pattern stays accepted for the life of the process, which also
spares the probe to every later apply, scenario step and keystroke; a refused
one is judged afresh each time, so one unlucky run does not stick.

The probe runs TWICE before refusing, and the second run decides, because a
wall clock cannot tell "this pattern burned five milliseconds" from "this
thread lost the CPU for five milliseconds". The obvious answer to that is a CPU
clock, and it does not work here: `time.get_clock_info("thread_time")` REPORTS
a resolution of 1e-07 on this platform and MEASURES 15.625 ms (200 000 reads
returned six distinct values, 2026-09-02), which cannot see a 5 ms budget at
all. A second run can: a scheduling hiccup does not repeat in the same place,
and backtracking does, every time, deterministically. A good pattern never
pays for this - it takes the first run only, at 0.04 ms.
"""
try:
# A user pattern like "[a-z[0-9]]" makes `re` emit a FutureWarning ("possible
# nested set"). It is not an error and the pattern still compiles - but the
# warning goes to stderr, which in a windowed build DOES NOT EXIST, and in the
# CLI lands in the middle of the log channel. Either way it is noise the user
# can do nothing about, so it is swallowed here; a pattern that is genuinely
# broken still raises re.error below.
# broken still raises below.
with warnings.catch_warnings():
warnings.simplefilter("ignore", FutureWarning)
warnings.simplefilter("ignore", DeprecationWarning)
rx = re.compile(pattern, re.IGNORECASE)
except re.error as exc:
raise _err("errors.bad_filter_regex", field, term) from exc
_refuse_if_too_slow(rx, field, term)
# Not only re.error: `a{99999999999}` raises OverflowError and a few thousand
# nested groups RecursionError, and neither is a ValueError - the one thing every
# caller catches. MEASURED 2026-09-28: `--dry-run` exited 1 instead of CONFIG, the
# expression tester raised although it promises it never does, and the Control
# page raised on every keystroke. This is the single place all of them pass.
except (re.error, OverflowError, RecursionError) as exc:
raise _Refused("errors.bad_filter_regex") from exc
if _blows_the_budget(rx) and _blows_the_budget(rx):
raise _Refused("errors.filter_regex_too_slow")
return rx


def _compile_regex(pattern, field, term):
pattern = pattern.strip()
if not pattern:
raise _err("errors.bad_filter_regex", field, term)
try:
return _accepted_regex(pattern)
except _Refused as refused:
raise _err(refused.key, field, term) from refused


def _is_glob(body):
return "*" in body or "?" in body

Expand Down
11 changes: 6 additions & 5 deletions beantester/nettools/exprtest.py
Original file line number Diff line number Diff line change
Expand Up @@ -30,11 +30,12 @@
BAD_EXPRESSION = "bad_expression"

# The longest test value taken. The expression parser vets a regular expression by
# timing it on probes of up to 45 characters (`matchers._REGEX_PROBE_LENGTHS`), and
# runs on the UI thread, as the Control page's live validation does - so a value far
# longer than anything the vetting saw is refused rather than run through a
# pattern nobody has timed on it. 256 is far past any process name this tool has
# met; an address or a port never gets near it.
# timing it on probes of up to 45 characters (`matchers._REGEX_PROBE_LENGTHS`), each
# also followed by a character it cannot consume, and this runs on the UI thread.
# Not cut down to 45: MEASURED 2026-09-28, the shapes that still pass that vetting
# cost at most ~1 ms per search at 256 characters (`(\w|\d)+z`), while 45 would
# refuse long process names the tester exists to try. 256 is far past any process
# name this tool has met; an address or a port never gets near it.
MAX_VALUE_CHARS = 256

# ASCII digits, not ``str.isdigit()``: MEASURED 2026-09-23, for "443" written in
Expand Down
29 changes: 18 additions & 11 deletions beantester/settings.py
Original file line number Diff line number Diff line change
Expand Up @@ -440,6 +440,8 @@ def apply_targeting(engine, target, log=lambda *_: None, announce=True):

Returns the live :class:`~beantester.targeting.ProcessTargeting` (iterable,
``len()``-able), or ``None`` when targeting is off / could not be resolved.
An expression that cannot be read changes nothing: the engine keeps the target
it had, and that is what is returned.
The object keeps re-resolving itself while the session runs, so a connection
the target opens a second from now is impaired too - the old code handed the
engine a frozen set of ports and everything opened afterwards escaped it.
Expand All @@ -453,9 +455,11 @@ def apply_targeting(engine, target, log=lambda *_: None, announce=True):
try:
matcher = parse_matcher(expr, KIND_PROCESS, TARGET_FIELD)
except ValueError as e:
# Left as it was, like the destination in apply_settings: switching
# targeting off here means impairing every connection in the filter,
# which is the widest answer to an expression that could not be read.
log(f"{T('log.targeting_error')}: {e}")
engine.set_target(False)
return None
return engine.targeting()
if matcher.is_empty:
engine.set_target(False)
return None
Expand Down Expand Up @@ -608,11 +612,14 @@ def apply_settings(engine, s, log=lambda *_: None):
try:
dest = (bool(dst_ip or dst_port), *compile_endpoint(dst_ip, dst_port))
except ValueError as e:
# Tolerant like the schedule below: a bad expression disables destination
# targeting instead of killing a scenario thread. The GUI and the CLI
# validate up front (validate_settings), so a user never reaches this.
# Tolerant like the schedule below - a scenario thread must not die of it
# - but NEVER by switching the destination off: off means "impair
# everything", so a field that could not be read widened the session to
# the whole machine (external review, P3-13). It is left as it was. The
# GUI and the CLI validate up front, and an accepted regex stays accepted
# (matchers._accepted_regex), so a user does not reach this.
log(f"{T('log.filter_skipped')}: {e}")
dest = (False, *compile_endpoint(None, None))
dest = None
# The same shape as the pair below, and said for the same reason: two "only"
# switches that exclude each other leave nothing to aim at, the symptom is a
# session that changes nothing, and that looks like a broken tool rather than
Expand All @@ -636,15 +643,14 @@ def apply_settings(engine, s, log=lambda *_: None):
log(T("log.asym_one_way_filter"))
block_ip = setting_expression("block_ip", g("block_ip"))
block_port = setting_expression("block_port", g("block_port"))
block = None # None = leave the block alone
try:
block = (bool(block_ip or block_port), *compile_endpoint(block_ip, block_port),
bool(g("block_reject")))
except ValueError as e:
# Tolerant like destination above: a bad expression disables blocking
# instead of killing a scenario thread. GUI and CLI validate up front.
# The mode goes with it: with no block there is nothing to refuse.
# Tolerant like destination above, and by the same rule: a field that could
# not be read is left as it was, mode included, rather than guessed at.
log(f"{T('log.filter_skipped')}: {e}")
block = (False, *compile_endpoint(None, None), False)
try:
schedule = parse_schedule(g("rate_schedule"))
except ValueError as e:
Expand All @@ -668,7 +674,8 @@ def apply_settings(engine, s, log=lambda *_: None):
engine.set_ip_family(bool(g("ipv4_only")), bool(g("ipv6_only")))
engine.set_lan(bool(g("lan_mode")))
engine.set_internet_only(bool(g("internet_only")))
engine.set_block(*block)
if block is not None:
engine.set_block(*block)
engine.set_advanced(g("syn_drop"), g("max_size"))
engine.set_spike(g("spike_prob"), g("spike_ms"))
engine.set_nat(g("nat_timeout"))
Expand Down
2 changes: 1 addition & 1 deletion lang/en.json
Original file line number Diff line number Diff line change
Expand Up @@ -296,7 +296,7 @@
"log.duration_reached": "Time limit reached ({v} s) - session stopped.",
"log.engine_fault": "Engine fault: {e} - the session was stopped, your network is back to normal.",
"log.error": "Error",
"log.filter_skipped": "This expression could not be read, so it was switched off for this session",
"log.filter_skipped": "This expression could not be read, so this field was left as it was",
"log.ipv4_and_ipv6_only": "IPv4 addresses only and IPv6 addresses only are both on. No packet is both, so nothing will be impaired.",
"log.lan_and_internet_only": "LAN mode and Internet only are both on - nothing but loopback gets through.",
"log.layout_reset": "Window layout reset.",
Expand Down
2 changes: 1 addition & 1 deletion lang/pl.json
Original file line number Diff line number Diff line change
Expand Up @@ -296,7 +296,7 @@
"log.duration_reached": "Osiągnięto limit czasu ({v} s) - sesja zatrzymana.",
"log.engine_fault": "Awaria silnika: {e} - sesja została zatrzymana, sieć działa normalnie.",
"log.error": "Błąd",
"log.filter_skipped": "Nie udało się odczytać tego wyrażenia, więc zostało wyłączone na tę sesję",
"log.filter_skipped": "Nie udało się odczytać tego wyrażenia, więc to pole zostało bez zmian",
"log.ipv4_and_ipv6_only": "Włączone są naraz Tylko adresy IPv4 i Tylko adresy IPv6. Żaden pakiet nie jest jednym i drugim, więc nic nie zostanie zmienione.",
"log.lan_and_internet_only": "Tryb LAN i Tylko internet są włączone naraz - poza loopbackiem nic nie przejdzie.",
"log.layout_reset": "Układ okna zresetowany.",
Expand Down
2 changes: 1 addition & 1 deletion lang/zh.json
Original file line number Diff line number Diff line change
Expand Up @@ -296,7 +296,7 @@
"log.duration_reached": "已达到时间限制({v} 秒),会话已停止。",
"log.engine_fault": "引擎故障:{e}。会话已停止,网络已恢复正常。",
"log.error": "错误",
"log.filter_skipped": "无法解析此表达式,本次会话已将其关闭",
"log.filter_skipped": "无法解析此表达式,因此该字段保持不变",
"log.ipv4_and_ipv6_only": "同时启用了“仅 IPv4 地址”和“仅 IPv6 地址”。没有数据包同时属于两者,因此不会有任何改动。",
"log.lan_and_internet_only": "“局域网模式”和“仅互联网”同时启用,因此除环回流量外,其他流量都无法通过。",
"log.layout_reset": "窗口布局已重置。",
Expand Down
3 changes: 3 additions & 0 deletions tests/test_cli_runtime.py
Original file line number Diff line number Diff line change
Expand Up @@ -87,6 +87,9 @@ def test_exit_code_config_for_bad_input():
cases = {
"unknown preset": ["--preset", "nope", "--simulate"],
"bad expression": ["--dst-port", "80,abc", "--simulate"],
# OverflowError out of `re` used to escape as exit 1 with a traceback
"regex re cannot build": ["--dst-ip", "re:a{99999999999}", "--simulate"],
"regex too slow per packet": ["--dst-ip", r"re:^([\d:]+)+$", "--simulate"],
"bad schedule": ["--rate-schedule", "1:x:2", "--simulate"],
"out of range": ["--loss", "250", "--simulate"],
"negative duration": ["--duration", "-5", "--simulate"],
Expand Down
Loading
Loading