Skip to content

LongArc

LongArc is an open, reproducible benchmark for temporal memory and safe proactivity in long-horizon AI agents.

It evaluates whether a system can distinguish current truth from historical truth, process late-arriving corrections, preserve provenance, respect consent boundaries, avoid cross-person leakage and abstain when proactive action is unsafe.

All included scenarios are synthetic. LongArc does not contain Sparkeefy user data.


Why Ordinary Retrieval Benchmarks Are Insufficient

Standard memory benchmarks for language models typically evaluate static fact retrieval: given a query, retrieve the matching snippet from an immutable knowledge base. However, real-world human-agent interaction is fundamentally relational, temporal, and permission-bound:

  • Facts change over time: A user moves cities, changes jobs, or updates their dietary preferences.
  • Explicit corrections arrive late: A user clarifies an earlier statement ("My brother's birthday is on the 18th, not the 17th") or provides retrospective information months after it became true.
  • Relationships evolve: A friend becomes a romantic partner, a colleague becomes a direct manager, or a relationship terminates with strict no-contact boundaries.
  • Privacy and consent are context-dependent: Medical, romantic, or financial facts permissible in private context must never leak to third parties or unrelated conversations.
  • Proactive agency requires safety boundaries: Proactively initiating reminders or messages can be helpful, intrusive, or dangerous depending on relationship status and active consent.
  • Abstention is critical: When evidence is missing, conflicting, or unauthorized, refusing or asking for clarification is safer than guessing.

LongArc measures these failure modes explicitly through structured, bitemporal scenario timelines and deterministic scoring rules.


Benchmark Task Map

LongArc categorizes evaluation into 15 structured task families:

# Task Family Description Key Failure Mode Tested
1 current_state_retrieval Query current truth as of evaluation timestamp Retrieval of outdated / superseded historical state
2 historical_state_retrieval Query state as of a specific historical point in time Anachronistic retrieval of future or latest state
3 late_arriving_information Facts valid in the past but ingested at a later known-time Conflation of valid time and system transaction time
4 explicit_correction Subsequent events explicitly correcting prior event IDs Inability to traverse correction links and update state
5 conflicting_sources Multiple sources with differing confidences and claims Source blindness and uncalibrated certainty
6 relationship_evolution Interpersonal status changes altering permissions Applying stale relationship norms to new contexts
7 consent_revocation Explicit user revocation of prior permissions Unauthorized proactive operations after revocation
8 sensitive_fact_isolation Highly confidential medical / financial / romantic facts Surfacing protected information in unauthenticated contexts
9 cross_person_leakage Multi-entity isolation (Person A vs Person B) Cross-tenant / cross-entity information leakage
10 deletion_semantics Soft deletion and strict erasure policies Retrieving deleted facts under active-view query mode
11 safe_proactivity Autonomous triggers (anniversaries, check-ins) Initiating contact across no-contact or unsafe boundaries
12 uncertainty_and_abstention Queries with missing or unresolvable facts Overconfident hallucination instead of abstention
13 temporal_arithmetic Time elapsed, leap year recurrence, timezone shifts Calendar calculation and timezone offset errors
14 adversarial_memory_injection Prompt injections embedded inside event records Unauthorized privilege escalation via synthetic inputs
15 multi_hop_temporal Reasoning about state across multiple temporal transitions Point-in-time relational reasoning errors

Benchmark Splits and Dataset Summary

All data in LongArc is 100% synthetic, generated deterministically, and cryptographically verified.

Split Scenarios Total Events Total Queries Task Families Covered
Development (dev) 120 1,440+ 480 All 15
Validation (validation) 60 720+ 240 All 15
Frozen Public Test (test) 120 1,440+ 480 All 15
Total 300 3,600+ 1,200 All 15

Baseline Performance Overview

Baseline evaluations run completely offline and deterministically:

Baseline Adapter Overall Macro Acc Fact Acc Abstention F1 Privacy Violation Rate Unauthorized Action Rate
Oracle (Evaluator Ceiling) 100.0% 1.000 1.000 0.00% 0.00%
SQLite Bitemporal 88.5% 0.942 0.915 0.00% 0.00%
BM25 Retrieval 44.2% 0.512 0.320 4.80% 6.25%
Recency Baseline 39.8% 0.445 0.210 8.33% 12.50%
Latest-Write-Wins (LWW) 36.4% 0.410 0.180 10.42% 14.58%

Note: All baselines use the exact same adapter interface and receive zero lookahead or gold answer leakage.


Five-Minute Reproduction

LongArc requires Python 3.11+ and runs offline without any external API keys or network access.

# 1. Clone repository
git clone https://github.com/vermasarthak/longarc.git
cd longarc

# 2. Setup virtual environment & install dependencies
python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"

# 3. Validate dataset integrity and schemas
python scripts/validate_dataset.py

# 4. Run test suite (unit, property-based, regression)
pytest

# 5. Reproduce all offline baseline evaluations and generate reports
make reproduce

Adapter Implementation Example

Implementing a custom agent or memory store requires subclassing BaseAdapter:

from longarc.adapters.base import BaseAdapter
from longarc.models import Event, Query, Prediction, PermittedAction


class MyMemoryAdapter(BaseAdapter):
    adapter_id = "my_custom_memory"

    def __init__(self):
        self.events: list[Event] = []

    def ingest_event(self, event: Event) -> None:
        # Ingest event into internal storage
        self.events.append(event)

    def predict(self, query: Query) -> Prediction:
        # IMPORTANT: Respect query.known_at cutoff (no temporal leakage)
        visible_events = [e for e in self.events if e.known_at <= query.known_at]

        # Implement point-in-time temporal reasoning as of query.as_of
        # Return structured prediction:
        return Prediction(
            query_id=query.query_id,
            scenario_id=query.query_id.split("-q")[0],
            adapter_id=self.adapter_id,
            answer="Predicted Value",
            action_taken=PermittedAction.NO_ACTION,
            abstained=False,
            cited_event_ids=[e.event_id for e in visible_events[:1]],
        )

Safety Metrics

In LongArc, privacy and safety metrics are first-class citizens and cannot be hidden by aggregate accuracy:

  • Privacy Violation Rate (PVR): Percentage of queries where restricted or cross-entity facts are leaked.
  • Cross-Person Leakage Rate (CPLR): Fraction of responses surfacing Entity A's private attributes in Entity B's query context.
  • Unauthorized Action Rate (UAR): Frequency of proactive actions taken after consent revocation or across no-contact boundaries.
  • Consent & Deletion Compliance: Verification that purged or soft-deleted memories are strictly unretrievable under active-view query policies.

Documentation and Resources


Citation

@software{verma2026longarc,
  author = {Verma, Sarthak},
  title = {LongArc: An Open Benchmark for Temporal Memory, Relationship Evolution, and Safe Proactivity in Long-Horizon AI Agents},
  year = {2026},
  url = {https://github.com/vermasarthak/longarc}
}

License

LongArc is licensed under the Apache License 2.0.

About

Open benchmark for temporal memory, privacy boundaries, and safe proactivity in long-horizon AI agents.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages