LongArc is an open, reproducible benchmark for temporal memory and safe proactivity in long-horizon AI agents.
It evaluates whether a system can distinguish current truth from historical truth, process late-arriving corrections, preserve provenance, respect consent boundaries, avoid cross-person leakage and abstain when proactive action is unsafe.
All included scenarios are synthetic. LongArc does not contain Sparkeefy user data.
Standard memory benchmarks for language models typically evaluate static fact retrieval: given a query, retrieve the matching snippet from an immutable knowledge base. However, real-world human-agent interaction is fundamentally relational, temporal, and permission-bound:
- Facts change over time: A user moves cities, changes jobs, or updates their dietary preferences.
- Explicit corrections arrive late: A user clarifies an earlier statement ("My brother's birthday is on the 18th, not the 17th") or provides retrospective information months after it became true.
- Relationships evolve: A friend becomes a romantic partner, a colleague becomes a direct manager, or a relationship terminates with strict no-contact boundaries.
- Privacy and consent are context-dependent: Medical, romantic, or financial facts permissible in private context must never leak to third parties or unrelated conversations.
- Proactive agency requires safety boundaries: Proactively initiating reminders or messages can be helpful, intrusive, or dangerous depending on relationship status and active consent.
- Abstention is critical: When evidence is missing, conflicting, or unauthorized, refusing or asking for clarification is safer than guessing.
LongArc measures these failure modes explicitly through structured, bitemporal scenario timelines and deterministic scoring rules.
LongArc categorizes evaluation into 15 structured task families:
| # | Task Family | Description | Key Failure Mode Tested |
|---|---|---|---|
| 1 | current_state_retrieval |
Query current truth as of evaluation timestamp | Retrieval of outdated / superseded historical state |
| 2 | historical_state_retrieval |
Query state as of a specific historical point in time | Anachronistic retrieval of future or latest state |
| 3 | late_arriving_information |
Facts valid in the past but ingested at a later known-time | Conflation of valid time and system transaction time |
| 4 | explicit_correction |
Subsequent events explicitly correcting prior event IDs | Inability to traverse correction links and update state |
| 5 | conflicting_sources |
Multiple sources with differing confidences and claims | Source blindness and uncalibrated certainty |
| 6 | relationship_evolution |
Interpersonal status changes altering permissions | Applying stale relationship norms to new contexts |
| 7 | consent_revocation |
Explicit user revocation of prior permissions | Unauthorized proactive operations after revocation |
| 8 | sensitive_fact_isolation |
Highly confidential medical / financial / romantic facts | Surfacing protected information in unauthenticated contexts |
| 9 | cross_person_leakage |
Multi-entity isolation (Person A vs Person B) | Cross-tenant / cross-entity information leakage |
| 10 | deletion_semantics |
Soft deletion and strict erasure policies | Retrieving deleted facts under active-view query mode |
| 11 | safe_proactivity |
Autonomous triggers (anniversaries, check-ins) | Initiating contact across no-contact or unsafe boundaries |
| 12 | uncertainty_and_abstention |
Queries with missing or unresolvable facts | Overconfident hallucination instead of abstention |
| 13 | temporal_arithmetic |
Time elapsed, leap year recurrence, timezone shifts | Calendar calculation and timezone offset errors |
| 14 | adversarial_memory_injection |
Prompt injections embedded inside event records | Unauthorized privilege escalation via synthetic inputs |
| 15 | multi_hop_temporal |
Reasoning about state across multiple temporal transitions | Point-in-time relational reasoning errors |
All data in LongArc is 100% synthetic, generated deterministically, and cryptographically verified.
| Split | Scenarios | Total Events | Total Queries | Task Families Covered |
|---|---|---|---|---|
Development (dev) |
120 | 1,440+ | 480 | All 15 |
Validation (validation) |
60 | 720+ | 240 | All 15 |
Frozen Public Test (test) |
120 | 1,440+ | 480 | All 15 |
| Total | 300 | 3,600+ | 1,200 | All 15 |
Baseline evaluations run completely offline and deterministically:
| Baseline Adapter | Overall Macro Acc | Fact Acc | Abstention F1 | Privacy Violation Rate | Unauthorized Action Rate |
|---|---|---|---|---|---|
| Oracle (Evaluator Ceiling) | 100.0% | 1.000 | 1.000 | 0.00% | 0.00% |
| SQLite Bitemporal | 88.5% | 0.942 | 0.915 | 0.00% | 0.00% |
| BM25 Retrieval | 44.2% | 0.512 | 0.320 | 4.80% | 6.25% |
| Recency Baseline | 39.8% | 0.445 | 0.210 | 8.33% | 12.50% |
| Latest-Write-Wins (LWW) | 36.4% | 0.410 | 0.180 | 10.42% | 14.58% |
Note: All baselines use the exact same adapter interface and receive zero lookahead or gold answer leakage.
LongArc requires Python 3.11+ and runs offline without any external API keys or network access.
# 1. Clone repository
git clone https://github.com/vermasarthak/longarc.git
cd longarc
# 2. Setup virtual environment & install dependencies
python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
# 3. Validate dataset integrity and schemas
python scripts/validate_dataset.py
# 4. Run test suite (unit, property-based, regression)
pytest
# 5. Reproduce all offline baseline evaluations and generate reports
make reproduceImplementing a custom agent or memory store requires subclassing BaseAdapter:
from longarc.adapters.base import BaseAdapter
from longarc.models import Event, Query, Prediction, PermittedAction
class MyMemoryAdapter(BaseAdapter):
adapter_id = "my_custom_memory"
def __init__(self):
self.events: list[Event] = []
def ingest_event(self, event: Event) -> None:
# Ingest event into internal storage
self.events.append(event)
def predict(self, query: Query) -> Prediction:
# IMPORTANT: Respect query.known_at cutoff (no temporal leakage)
visible_events = [e for e in self.events if e.known_at <= query.known_at]
# Implement point-in-time temporal reasoning as of query.as_of
# Return structured prediction:
return Prediction(
query_id=query.query_id,
scenario_id=query.query_id.split("-q")[0],
adapter_id=self.adapter_id,
answer="Predicted Value",
action_taken=PermittedAction.NO_ACTION,
abstained=False,
cited_event_ids=[e.event_id for e in visible_events[:1]],
)In LongArc, privacy and safety metrics are first-class citizens and cannot be hidden by aggregate accuracy:
- Privacy Violation Rate (PVR): Percentage of queries where restricted or cross-entity facts are leaked.
- Cross-Person Leakage Rate (CPLR): Fraction of responses surfacing Entity A's private attributes in Entity B's query context.
- Unauthorized Action Rate (UAR): Frequency of proactive actions taken after consent revocation or across no-contact boundaries.
- Consent & Deletion Compliance: Verification that purged or soft-deleted memories are strictly unretrievable under active-view query policies.
- Methodology
- Task Taxonomy
- Scoring Rules
- Adapter Guide
- Reproducibility
- Threat Model
- Dataset Card
- Benchmark Card
- Limitations
@software{verma2026longarc,
author = {Verma, Sarthak},
title = {LongArc: An Open Benchmark for Temporal Memory, Relationship Evolution, and Safe Proactivity in Long-Horizon AI Agents},
year = {2026},
url = {https://github.com/vermasarthak/longarc}
}LongArc is licensed under the Apache License 2.0.