Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
185 changes: 185 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,185 @@
# CodeGen Compare Tool

Tool for comparing AUTOSAR MATLAB/Simulink codegen folders and identifying real changes while filtering generator noise.

Repo: `codegen-compare-tool`
Package: `compare_tool`
Architecture: `docs/architecture.md`

This file contains **project invariants only**. Do not duplicate general coding guidance.

---

## Core Rule

**If a difference cannot be proven to be noise, it is a real change.**

Never hide, downgrade or silently ignore a potentially real change.

A scan/read/compare failure is an `error`, never an empty result.

---

## Verdicts

Supported file verdicts:

```text
identical
comment-only
ignorable-only
real-change
added
deleted
error
```

* Verdict logic has a single source of truth.
* Only noise verdicts may be folded/hidden by the UI.
* `real-change`, `added`, `deleted` and `error` must remain visible.
* UI filtering must never change the underlying verdict or counts.
* Reports and summaries must use the **raw scan result**, not the filtered UI state.

---

## Noise Rules

A noise rule must be conservative.

Every new rule should prove both:

1. The pattern alone is noise.
2. The same pattern next to a real change remains a real change.

Prefer text-based, anchored rules that preserve line structure where possible.

Never add a rule simply because it makes a diff smaller.

---

## Architecture

Keep shared decisions in one place.

* Shared comparison/view-model logic must not be duplicated between CLI, viewer and report.
* Shared visual roles belong in `theme.py`; do not add raw colour literals elsewhere.
* The comparison core must remain independent of Qt.
* `compare_tool/qtviewer/` is the only Qt-dependent area.
* Keep reusable parsing/model modules Qt-free so headless tests continue to work.

See `docs/architecture.md` before making structural changes.

---

## Dependencies

The comparison core and `compare_tool.pyz` must remain **Python standard-library only**.

PySide6 is allowed only for the desktop viewer and must be imported lazily.

The CLI must remain usable on locked-down machines with no installed third-party dependencies.

Python support:

```text
>= 3.8
```

Do not introduce syntax/runtime features unavailable on Python 3.8.

The HTML report must remain fully self-contained:

* CSS inline
* JavaScript inline
* No CDN
* No network requests

---

## Fail Safe

Prefer graceful degradation for non-critical UI resources.

Examples:

* Missing icons → keep the text label.
* Missing PySide6 → provide a clear viewer-install message.
* Legacy Windows console → handle encoding safely.

But **never degrade silently for comparison failures**.

A failed or incomplete comparison must be obvious and must not look like a clean comparison.

---

## Exit Codes

```text
0 No real changes
1 Real changes found
2 Compare incomplete / error
```

Exit code `2` is part of the CI contract and must never be suppressed by `--exit-zero`.

---

## Release

For a normal release:

1. Bump `compare_tool.__version__`.
2. Move the relevant CHANGELOG entries from `[Unreleased]` to the new version/date.
3. Let the release workflow build and publish.
4. Never overwrite an already published version.

`packaging/release_check.py` owns release preconditions.

---

## Documentation

The English documentation is the source of truth.

Vietnamese documentation in `docs/vi/` should translate meaning, not terminology.

Keep industry/tool terms in English where appropriate:

```text
port, runnable, calibration, noise, hunk, verdict, SWC, A2L, ARXML
```

When changing a documented behavior or claim, update the relevant documentation in the same change.

---

## Common Commands

```bash
# Compare
python -m compare_tool <old_gen> <new_gen> --report out.html

# Tests
python -m unittest discover -s tests

# Lint
python -m ruff check .

# Build EXE
.\build.ps1

# Build EXE + zipapp
.\build.ps1 -Pyz

# Zipapp only
.\build.ps1 -PyzOnly
```

Before committing:

```bash
python -m unittest discover -s tests
python -m ruff check .
```

Qt tests may skip when PySide6 is unavailable, so a green test run without PySide6 does not prove viewer tests were executed.
12 changes: 11 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,17 @@
All notable changes to this project are documented here. Versions follow
[semantic versioning](https://semver.org/).

## [Unreleased]
## [1.15.0]

### Added

- Add terminal-only comparisons with `--no-report`, including a complete file tree, AUTOSAR/A2L summaries and consistency warnings.

### Fixed

- Fix missed changes to string contents and Python/YAML indentation.
- Fix external function changes and reordered assignments with side effects being reported as noise.
- Fix incomplete comparisons when an added or deleted text file has invalid encoding.

## [1.14.1] — 2026-09-13

Expand Down
11 changes: 10 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -67,6 +67,15 @@ ZIP files can be compared directly:
python -m compare_tool baseline.zip current.zip --report report.html
```

### Print a terminal summary without a report

```bash
python -m compare_tool old_gen_folder new_gen_folder --no-report
```

Prints a folder tree with every file's verdict, AUTOSAR/A2L changes, and the same
consistency warnings as the HTML report. No code diff or HTML file is generated.

### Open the desktop viewer

```bash
Expand Down Expand Up @@ -105,7 +114,7 @@ Both use the **same comparison engine**, so they always produce the same compari
| ---------- | ------------------ | ------------------- |
| Best for | Interactive review | CI / automation |
| Input | Folders / ZIP | Folders / ZIP |
| Output | Interactive diff | HTML / JSON / SARIF |
| Output | Interactive diff | Terminal summary / HTML / JSON / SARIF |
| Build gate | — | Exit code |

---
Expand Down
4 changes: 2 additions & 2 deletions compare_tool/a2l_rules.py
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@

import re

from .c_rules import collapse_ws
from .langspec import SPECS, normalize_ws

# keywords are uppercase per the ASAM grammar, but /begin casing varies in
# the wild, so matching is case-insensitive and the kind is normalized
Expand Down Expand Up @@ -86,7 +86,7 @@ def strip_a2l_comments(text):

def a2l_shadow(text):
"""Normalized shadow for A2L: comments stripped + whitespace collapsed."""
return collapse_ws(strip_a2l_comments(text))
return normalize_ws(strip_a2l_comments(text), SPECS['a2l'])


def extract_objects(text):
Expand Down
36 changes: 30 additions & 6 deletions compare_tool/c_rules.py
Original file line number Diff line number Diff line change
Expand Up @@ -101,9 +101,10 @@ def strip_c_comments(text):
return ''.join(out)


def collapse_ws(text):
"""Collapse each line's whitespace runs to single spaces, strip edges."""
return '\n'.join(' '.join(line.split()) for line in text.split('\n'))
def collapse_ws(text: str) -> str:
"""Collapse C layout whitespace while preserving string/character data."""
from .langspec import SPECS, normalize_ws
return normalize_ws(text, SPECS['c'])


def c_shadow(text):
Expand Down Expand Up @@ -227,6 +228,7 @@ def detect_renames(old_shadow, new_shadow, hunks=None):
return None
old_ids = set(t for t in tokenize(old_shadow) if is_identifier(t))
new_ids = set(t for t in tokenize(new_shadow) if is_identifier(t))
calls = _call_names(old_shadow) | _call_names(new_shadow)
mapping = {}
for o, ns in fwd.items():
if len(ns) != 1:
Expand All @@ -238,6 +240,8 @@ def detect_renames(old_shadow, new_shadow, hunks=None):
continue # not a true rename (name still in use / swap)
if _is_authored_name(o) or _is_authored_name(n):
continue # RTE port / DWork field / macro-enum -- not this file's to rename
if (o in calls or n in calls) and not _generated_call_pair(o, n):
continue # a different callee is not a local variable rename
if not same_checksummed_object(o, n):
continue # consistent, but it names a different generated object
mapping[o] = n
Expand All @@ -248,7 +252,18 @@ def apply_rename_map(text, mapping):
"""Apply an identifier rename map to text (word-boundary safe)."""
if not mapping:
return text
return re.sub(r'[A-Za-z_]\w*', lambda m: mapping.get(m.group(0), m.group(0)), text)
return TOKEN_RE.sub(lambda m: mapping.get(m.group(0), m.group(0)), text)


def _call_names(text: str) -> set:
tokens = tokenize(text)
return {a for a, b in zip(tokens, tokens[1:]) if b == '(' and is_identifier(a)}


def _generated_call_pair(old: str, new: str) -> bool:
# A generic suffix (get_x/get_y) alone cannot identify a generated callee.
root = checksum_root(old)
return root is not None and root == checksum_root(new)


# --- MATLAB codegen autogenerated-name noise (Simulink/Embedded Coder) ---
Expand Down Expand Up @@ -385,13 +400,15 @@ def canonical_generated(text):
"""
def sub(m):
tok = m.group(0)
if not is_identifier(tok):
return tok
root = generated_root(tok)
if root is not None:
return root
root = checksum_root(tok)
return root if root is not None else tok

return re.sub(r'[A-Za-z_]\w*', sub, text)
return TOKEN_RE.sub(sub, text)


def _map_has_cycle(mapping):
Expand All @@ -418,7 +435,12 @@ def autogen_noise_map(old_lines, new_lines, old_ids=None, new_ids=None):
collected, or when the map contains a cycle (i_0 <-> i_1 index swap /
rotation is a real semantic change).
"""
accept = lambda a, b: is_autogen_name_pair(a, b, old_ids, new_ids) # noqa: E731
calls = _call_names('\n'.join(old_lines + new_lines))

def accept(a, b):
if (a in calls or b in calls) and not _generated_call_pair(a, b):
return False
return is_autogen_name_pair(a, b, old_ids, new_ids)
if len(old_lines) != len(new_lines):
# A shorter generated name lets an argument fit on one line where it
# used to wrap at 80 columns, so the two sides hold the same statements
Expand Down Expand Up @@ -627,6 +649,8 @@ def _parse_scalar_stmt(line):
return None # '==' comparison, or an empty RHS -- not an assignment
if _CALL_RE.search(rhs):
return None # a call may have side effects; moving it is not safe
if re.search(r'\+\+|--|(?:<<|>>|[+*/%&^|\-])=|(?<![=!<>])=(?!=)', rhs):
return None # RHS writes are not represented by the scalar read set
if lhs in C_KEYWORDS:
return None
content = canonical_generated(body)
Expand Down
8 changes: 5 additions & 3 deletions compare_tool/diff_engine.py
Original file line number Diff line number Diff line change
Expand Up @@ -265,7 +265,9 @@ def _build_variants(old_text, new_text, ruleset, rename_map, ext='', user_rules=
kind wins the label when both explain a hunk. A hunk only a user rule
explains is labelled with that rule's name and is ignorable-only, never
comment-only -- comment is a built-in category with its own report rules."""
cw = c_rules.collapse_ws
def cw(text):
return langspec.normalize_ws(text, langspec.SPECS.get(ruleset,
langspec.SPECS['c']))
variants = [('whitespace', _lines(cw(old_text)), _lines(cw(new_text)))]
if ruleset == 'c':
old_nc = c_rules.strip_c_comments(old_text)
Expand Down Expand Up @@ -308,8 +310,8 @@ def _build_variants(old_text, new_text, ruleset, rename_map, ext='', user_rules=
# the pass-2 diff runs on, so variant and shadow stay in sync.
spec = langspec.SPECS[ruleset]
variants.append(('comment',
_lines(cw(langspec.strip_comments(old_text, spec))),
_lines(cw(langspec.strip_comments(new_text, spec)))))
_lines(langspec.shadow(old_text, spec)),
_lines(langspec.shadow(new_text, spec))))
# user rules last: a hunk a built-in kind already explains keeps that kind
for r in user_rules:
if r.applies_to(ext):
Expand Down
Loading