Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 3 additions & 9 deletions .github/workflows/docs.yml
Original file line number Diff line number Diff line change
Expand Up @@ -19,16 +19,10 @@ jobs:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: '3.12'
- uses: astral-sh/setup-uv@v5

# ADJUST THIS: install all dependencies (including pdoc)
- run: pip install -e .
- run: pip install pdoc
# ADJUST THIS: build your documentation into docs/.
# We use a custom build script for pdoc itself, ideally you just run `pdoc -o docs/ ...` here.
- run: pdoc src/ultk -d google --math -o ./docs
- run: uv sync --group dev
- run: uv run pdoc src/ultk -d google --math -o ./docs
Comment thread
shanest marked this conversation as resolved.

- uses: actions/upload-pages-artifact@v3
with:
Expand Down
11 changes: 4 additions & 7 deletions .github/workflows/pypi-publish.yml
Original file line number Diff line number Diff line change
Expand Up @@ -13,15 +13,12 @@ jobs:
id-token: write # IMPORTANT: this permission is mandatory for trusted publishing
steps:
# retrieve your distributions here
- uses: actions/checkout@v3
- uses: actions/setup-python@v4
- uses: actions/checkout@v4
- uses: astral-sh/setup-uv@v5
Comment thread
shanest marked this conversation as resolved.
with:
python-version: '3.11'
- name: Install package
run: pip install --upgrade build
pip install -e .
python-version: "3.13"
- name: Build dist
run: python -m build
run: uv build
- name: Publish package distributions to PyPI
uses: pypa/gh-action-pypi-publish@release/v1
with:
Expand Down
15 changes: 6 additions & 9 deletions .github/workflows/test.yml
Original file line number Diff line number Diff line change
Expand Up @@ -6,14 +6,11 @@ jobs:
build:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v3
- name: Set up Python
uses: actions/setup-python@v4
- uses: actions/checkout@v4
- uses: astral-sh/setup-uv@v5
Comment thread
shanest marked this conversation as resolved.
with:
python-version: '3.11'
- name: Install package
run: pip install -e .
python-version: "3.13"
- name: Install dependencies
run: uv sync --group dev
- name: Test with pytest
run: |
pip install pytest
pytest
run: uv run pytest src/tests/
64 changes: 64 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
# CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

## Project Overview

ULTK (Unnatural Language ToolKit) is a Python library for computational semantic typology research — specifically for "efficient communication" analyses that explain natural language structure in terms of competing pressures: minimizing cognitive complexity vs. maximizing communicative accuracy.

## Commands

```bash
# Install all dependencies (including dev group for tests)
uv sync --group dev

# Run all tests
uv run pytest src/tests/

# Run a single test file
uv run pytest src/tests/test_language.py

# Run a single test by name
uv run pytest src/tests/test_language.py::TestLanguage::test_name

# Format code (Black is enforced via CI on PRs)
black src/
Comment thread
shanest marked this conversation as resolved.
```

Tests are discovered automatically by pytest from `src/tests/`. The CI workflow runs `uv run pytest src/tests/` from the repo root.

## Architecture

### Two Main Modules

**`ultk.language`** — Core data structures for semantic representations:
- `semantics.py`: `Referent` (immutable semantic object), `Universe` (collection of Referents with a prior distribution), `Meaning` (mapping from Universe to arbitrary type T — e.g., booleans for truth values)
- `language.py`: `Expression` (form + meaning pair), `Language` (frozenset of Expressions sharing a Universe). Helper `aggregate_expression_complexity()` bridges language and effcomm.
- `sampling.py`: Generators for all meanings, expressions, and languages from a universe — used to enumerate the full hypothesis space.
- `grammar/`: A probabilistic context-free grammar (PCFG) framework for building expressions as programs in a Language of Thought. `grammar.py` defines `Rule` and `Grammar`/`GrammaticalExpression`; `likelihood.py` provides scoring functions; `inference.py` handles MDL/Bayesian inference.

**`ultk.effcomm`** — Efficient communication analysis tools:
- `agent.py`: RSA (Rational Speech Act) agents — `LiteralSpeaker`, `LiteralListener`, `PragmaticSpeaker`, `PragmaticListener` — represented as weight matrices.
- `informativity.py`: `informativity()` and `communicative_success()` — compute how well a language supports communication (vectorized as `diag(prior) @ S @ R ⊙ U`).
- `tradeoff.py`: Pareto front computation (`pareto_optimal_languages`, `non_dominated_2d`, `dominates`) for simplicity/informativeness trade-off analysis.
- `optimization.py`: `EvolutionaryOptimizer` — iterative algorithm to approximate the Pareto frontier via mutations (`AddExpression`, `RemoveExpression`).
- `sampling.py`: `get_hypothetical_variants()` — generates null-hypothesis languages by permuting speaker weight matrices.
- `analysis.py`: Aggregation utilities for building results DataFrames.

**`ultk.util`**:
- `frozendict.py`: `FrozenDict` — an immutable dict used extensively as keys in frozen dataclasses.
- `io.py`: I/O helpers.

### Key Design Patterns

- Core objects (`Universe`, `Meaning`, `Expression`) are **frozen/immutable** (`@dataclass(frozen=True)` or manual `_frozen` flag), enabling hashing and use as dict keys.
- `Meaning` stores its mapping as a `tuple[T, ...]` indexed parallel to `Universe.referents`, with `_ref_to_idx` for O(1) lookup. Access via `meaning[referent]`.
- `Language` stores expressions as a `frozenset` — order-independent, hashable.
- Grammar rules are defined via Python type annotations; `Rule.from_callable()` introspects function signatures to build rules automatically.

### Examples

`src/examples/` contains complete worked analyses:
- `indefinites/` — efficient communication analysis of indefinite pronouns
- `modals/` — semantic universals for modals
- `learn_quant/` — quantifier learning
23 changes: 19 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,13 +9,21 @@ Read the [documentation](https://clmbr.shane.st/ultk).

## Installing ULTK

First, set up a virtual environment (e.g. via [miniconda](https://docs.conda.io/en/latest/miniconda.html), `conda create -n ultk python=3.11`, and `conda activate ultk`).
ULTK requires Python 3.13+. We recommend using [uv](https://docs.astral.sh/uv/) to manage dependencies.

1. Download or clone this repository and navigate to the root folder.

2. Install ULTK (We recommend doing this inside a virtual environment)
2. Install ULTK and all dependencies:

`pip install -e .`
```
uv sync
```

Alternatively, if you prefer pip inside an activated virtual environment:

```
pip install -e .
```

## Getting started

Expand All @@ -33,7 +41,14 @@ The source code is available on github [here](https://github.com/CLMBRs/ultk).

## Testing

Unit tests are written in [pytest](https://docs.pytest.org/en/7.3.x/) and executed via running `pytest` in the `src/tests` folder.
Unit tests are written in [pytest](https://docs.pytest.org/en/7.3.x/) and executed via:

```
uv sync --group dev
uv run pytest src/tests/
```
Comment thread
shanest marked this conversation as resolved.

Or, if inside an activated virtual environment: `pytest src/tests/`.

## References

Expand Down
9 changes: 5 additions & 4 deletions pyproject.toml
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
[project]
name = "ultk"
version = "0.0.1c"
version = "0.1.0"
authors = [
{ name="Chris Haberland", email="haberc@uw.edu"},
{ name="Nathaniel Imel", email="nimel@uci.edu"},
Expand All @@ -16,18 +16,19 @@ classifiers = [
]
license = {file = "LICENSE.txt"}
dependencies = [
"mypy",
"numpy",
"nltk",
"pyyaml",
"pandas",
"plotnine",
"pathos",
"pytest",
"scipy>=1.7.3",
"scipy-stubs[scipy]>=1.16.3.3",
"tqdm",
]

[dependency-groups]
dev = ["mypy", "pytest", "scipy-stubs[scipy]>=1.16.3.3", "pdoc"]

[project.urls]
"Homepage" = "https://clmbr.shane.st/ultk"
"Bug Tracker" = "https://github.com/CLMBRs/ultk/issues"
Expand Down
4 changes: 0 additions & 4 deletions setup.py

This file was deleted.

Empty file added src/tests/__init__.py
Empty file.
2 changes: 1 addition & 1 deletion src/tests/test_grammar.py
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ def test_meaning(self):
parsed_expression = TestGrammar.grammar.parse(TestGrammar.geq2_expr_str)
expr_meaning = parsed_expression.evaluate(TestGrammar.universe)
goal_meaning = Meaning(
{referent: referent.num > 2 for referent in TestGrammar.referents},
tuple(referent.num > 2 for referent in TestGrammar.referents),
TestGrammar.universe,
)
assert expr_meaning == goal_meaning
Expand Down
24 changes: 10 additions & 14 deletions src/tests/test_language.py
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,6 @@

from ultk.language.language import Expression, Language
from ultk.language.semantics import Referent, Universe, Meaning
from ultk.util.frozendict import FrozenDict


class TestLanguage:
Expand All @@ -27,35 +26,35 @@ class TestLanguage:
dog = Expression(
form="dog",
meaning=Meaning(
mapping=FrozenDict({ref: ref.name == "dog" for ref in uni_refs}),
mapping=tuple(ref.name == "dog" for ref in uni_refs),
universe=uni,
),
)
cat = Expression(
form="cat",
meaning=Meaning(
mapping=FrozenDict({ref: ref.name == "cat" for ref in uni_refs}),
mapping=tuple(ref.name == "cat" for ref in uni_refs),
universe=uni,
),
)
tree = Expression(
form="tree",
meaning=Meaning(
mapping=FrozenDict({ref: ref.name == "tree" for ref in uni_refs}),
mapping=tuple(ref.name == "tree" for ref in uni_refs),
universe=uni,
),
)
shroom = Expression(
form="shroom",
meaning=Meaning(
mapping=FrozenDict({ref: ref.name == "shroom" for ref in uni_refs}),
mapping=tuple(ref.name == "shroom" for ref in uni_refs),
universe=uni,
),
)
bird = Expression(
form="bird",
meaning=Meaning(
mapping=FrozenDict({ref: ref.name == "bird" for ref in uni_refs}),
mapping=tuple(ref.name == "bird" for ref in uni_refs),
universe=uni,
),
)
Expand All @@ -65,7 +64,7 @@ class TestLanguage:
lang_subset_expr = Language(expressions=tuple([dog, cat, tree]))
lang_of_different_order = Language(expressions=tuple([dog, cat, shroom, tree]))

def test_exp_subset(self):
def test_exp_can_express_positive(self):
assert TestLanguage.dog.can_express(Referent("dog", {"phylum": "animal"}))

def test_exp_subset(self):
Expand All @@ -83,11 +82,8 @@ def test_language_universe_check(self):
Expression(
form="dog",
meaning=Meaning(
mapping=FrozenDict(
{
ref: ref.name == "dog"
for ref in TestLanguage.uni.referents
}
mapping=tuple(
ref.name == "dog" for ref in TestLanguage.uni.referents
),
universe=TestLanguage.uni2,
),
Expand All @@ -98,8 +94,8 @@ def test_language_universe_check(self):
def test_language_degree(self):
def isAnimal(exp: Expression) -> bool:
print("checking phylum of " + str(exp))
for k, v in exp.meaning.mapping.items():
if v and k.phylum != "animal":
for ref, v in zip(exp.meaning.universe.referents, exp.meaning.mapping):
if v and ref.phylum != "animal":
return False
return True

Expand Down
18 changes: 14 additions & 4 deletions src/ultk/language/grammar/grammar.py
Original file line number Diff line number Diff line change
Expand Up @@ -100,7 +100,12 @@ def _and(p1: bool, p2: bool) -> bool:
rhs: tuple[Any, ...] | None = tuple(arg.annotation for arg in args.values())
# if one type annotation is class, a type of Referent, treat this as a terminal, no children = None RHS
# TODO: make this more general?
if rhs and len(rhs) == 1 and inspect.isclass(rhs[0]) and issubclass(rhs[0], Referent):
if (
rhs
and len(rhs) == 1
and inspect.isclass(rhs[0])
and issubclass(rhs[0], Referent)
):
rhs = None
return cls(
name=rule_name,
Expand Down Expand Up @@ -168,15 +173,20 @@ def complement(self) -> Meaning:
the expression evaluates to False."""

return Meaning(
tuple(set(self.meaning.universe.referents) - set(self.meaning.referents)),
tuple(not val for val in self.meaning.mapping),
self.meaning.universe,
)

def draw_referent(self, complement=False):
"""Get a random referent from the meaning's referents."""
universe_refs = self.meaning.universe.referents
if complement:
return random.choice(list(self.complement().referents))
return random.choice(list(self.meaning.referents))
return random.choice(
[r for r, v in zip(universe_refs, self.meaning.mapping) if not v]
)
return random.choice(
[r for r, v in zip(universe_refs, self.meaning.mapping) if v]
)

def to_dict(self) -> dict:
the_dict = super().to_dict()
Expand Down
3 changes: 1 addition & 2 deletions src/ultk/language/language.py
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,6 @@
from dataclasses import dataclass
from typing import Callable, Generic, Iterable, TypeVar
from ultk.language.semantics import Meaning, Referent, Universe
from ultk.util.frozendict import FrozenDict

# TODO: require Python 3.12 and use type parameter syntax instead? https://docs.python.org/3/reference/compound_stmts.html#type-params
T = TypeVar("T")
Expand All @@ -30,7 +29,7 @@ class Expression(Generic[T]):
# useful for hashing in certain cases
# (e.g. a GrammaticalExpression which has not yet been evaluate()'d and so does not yet have a Meaning)
form: str = ""
meaning: Meaning[T] = Meaning(FrozenDict(), Universe(tuple(), tuple()))
meaning: Meaning[T] = Meaning(tuple(), Universe(tuple(), tuple()))

def can_express(self, referent: Referent) -> bool:
"""Return True if the expression can express the input single meaning point and false otherwise."""
Expand Down
3 changes: 2 additions & 1 deletion src/ultk/language/sampling.py
Original file line number Diff line number Diff line change
Expand Up @@ -32,7 +32,8 @@ def all_meanings(universe: Universe) -> Generator[Meaning, None, None]:
"""Generate all Meanings (sets of Referents) from a given Universe."""
referents = universe.referents
for refset in powerset(referents):
yield Meaning(refset, universe)
refset_set = set(refset)
yield Meaning(tuple(ref in refset_set for ref in referents), universe)


def all_expressions(meanings: Iterable[Meaning]) -> Generator[Expression, None, None]:
Expand Down
Loading