Skip to content

Own address DB + cache improvements (reduce provider dependency) #2

Description

@BleedingDev

Goal

Reduce cost and third‑party dependency by building and using our own address dataset over time, while keeping suggestion quality high.

Who Uses This

  • Operators who want lower provider costs and fewer external dependencies.
  • Customers who want faster and more reliable suggestions.
  • Developers who need predictable behavior even when upstream providers degrade.

What “Done” Means (Business)

  • The system stores addresses observed during normal operation to build a reusable dataset.
  • Future suggestion requests can be served from our own dataset when possible, reducing external calls.
  • The product captures a signal for “accepted” addresses (what users actually select) to improve relevance and confidence.
  • The system still falls back to external providers when local coverage is insufficient.
  • Data handling has clear privacy and retention expectations (no user identification required to benefit from the dataset).

Constraints

  • Must not meaningfully slow down live suggestions.
  • Must support gradual rollout and safe behavior (never return obviously wrong merges).

Out of Scope

  • Bulk importing external address datasets.
  • Full postal validation/correction.
  • Multi-region/high-availability database replication.

Technical notes: #2 (comment)

Activity

  1. coderabbitai commented on Dec 28, 2025

    @coderabbitai

    📝 CodeRabbit Plan Mode

    Generate an implementation plan and prompts that you can use with your favorite coding agent.

    • Create Plan
    Examples

    🔗 Similar Issues

    Related Issues

    👤 Suggested Assignees

    🧪 Issue enrichment is currently in open beta.

    You can configure auto-planning by selecting labels in the issue_enrichment configuration.

    To disable automatic issue enrichment, add the following to your .coderabbit.yaml:

    issue_enrichment:
      auto_enrich:
        enabled: false

    💬 Have feedback or questions? Drop into our discord!

  2. BleedingDev commented on Dec 28, 2025

    @BleedingDev
    ContributorAuthor

    Technical Notes

    Current Foundations

    • SQLite DB + migrations: apps/service-bun/src/sqlite.ts
    • L1/L2 cache + SWR: apps/service-bun/src/cache.ts
    • Search logging: apps/service-bun/src/search-log.ts

    Proposed Data Model

    • Canonical address table (addresses) + provider-specific table (address_sources) so one canonical address can have multiple sources.
    • Conservative dedupe key (“canonical key”) computed from normalized address parts; allow null when data is insufficient.
    • Optional full-text search (SQLite FTS5) over labels + address parts to serve suggestions from local DB.

    Suggested Runtime Flow

    1. L1 cache
    2. L2 cache (SQLite)
    3. Local address DB search (FTS/prefix search)
    4. External provider(s) as fallback
    5. Persist newly observed addresses asynchronously (do not block response)

    Acceptance Signal

    • Add an acceptance endpoint (POST /accept) to record which suggestion was selected.
    • Increment acceptance counters to improve ranking/confidence for future results.

    Ownership + Reuse

    • Add shared types + normalization helpers in packages/core (e.g. computeCanonicalKey, canonical address record type) so other services can reuse the model.

    Safety/Privacy

    • No user identifiers required.
    • Treat session tokens as short-lived correlation only (store only where needed for analytics, not in the canonical address DB).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:cache-dbCaching / address databaseepicEpic (parent issue)priority:p2Post-MVP / later

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions