Skip to content

Recognize complete BibTeX block starts #583

Description

@claell

Problem

The splitter recognizes a block start with @\w* followed only by spaces or tabs before the opening delimiter. This is narrower than the BibTeX data-language structure in two important ways:

  • line endings and other ordinary whitespace may separate an entry type from its outer { or ( delimiter; and
  • entry-type identifiers may contain valid non-word punctuation such as -, :, or /.

Candidates outside the regular expression are currently retained as implicit comments rather than becoming parsed blocks or explicit failures. That makes a structural input disappear from the bibliography model without any parse failure.

Proposed behavior

Recognize the complete block-type token up to whitespace or a reserved BibTeX syntax character, locate the following standard outer delimiter across whitespace, and use the delimiter position rather than assuming it immediately follows the regex match.

Add a matrix requiring every recognized candidate to become either:

  • a parsed block for the currently supported curly-brace form; or
  • an explicit failed block for the currently unsupported parenthesis form.

This work builds on the explicit parenthesis-failure handling in #533 / PR #571 and exact special-command matching in #568 / PR #574. The implementation PR should retain those as separate prerequisite commits so the broader parser change remains reviewable.

Activity

  1. claell commented on Jul 16, 2026

    @claell
    ContributorAuthor

    A validated draft implementation is available in PR #584: #584. The first two commits preserve the explicit dependencies on PRs #571 and #574; the final two commits are the new issue #583 behavior and visibility matrix.

  2. MiWeiss commented on Sep 2, 2026

    @MiWeiss
    Collaborator

    Hi @claell — I'm posting this same note on all of your July 16 issues and PRs (#567–#597), so apologies for the form-letter feel.

    I'm closing this. The batch — 13 issues, 18 PRs, ~3,000 lines, opened within a two-hour window with spec-style language and forward-references to PR numbers that didn't exist yet — reads as AI-generated rather than something you hit and verified by hand. Reviewing it properly would cost me more time than just fixing the parser myself.

    If this is a real bug you've actually hit: reopen it with a concrete repro and a small, human-verified fix, and I'll review it. Any nontrivial design or API choice should be discussed in the issue first, before code is written. Otherwise, please disclose and verify AI-assisted contributions before submitting them in future.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions