VSE Text Based Editing is a Blender 5.2 LTS extension for editing complete Movie or audio-only base channels as text documents. Each supported strip on a channel becomes a media item; real spaces between strips remain timeline gaps. The implementation and project format are Blender-native.
Warning
Active development: This is an early 0.x release. Workflows, project data and UI may change between versions. Version 0.3 replaces the Sequencer overlay with an area-local mode in Blender's native Text Editor. Version 0.3.1 fixes activation and update cleanup while Blender exposes its restricted registration context. Version 0.3.2 moves the full Text Based Editor switch beside the Text datablock Open control and adds playhead word highlighting. Version 0.3.3 applies document cuts incrementally and keeps the playhead anchored to the edited spoken content. Version 0.3.8 expands removed-word cuts safely by 80 ms per edge, adds project-wide and per-word timing controls without changing the transcript, restores normal Space play/pause behavior in the Text Based Editor, and adds reversible batch pause removal with per-project detection controls. Native Text Editor selection now takes visual priority over removed-word decoration, and the pause menu can restore the configured add-on defaults. Version 0.4.0 adds non-destructive linear audio smoothing. Generated Sound strips overlap on automatically allocated free channels while video timing, the logical EDL and project duration remain unchanged. Version 0.4.1 keeps generated audio on its source channel whenever possible, confines AV audio below its Movie channel and natively connects every generated Movie/Sound segment pair. Version 0.4.2 leaves the VSE timeline completely unchanged after initial transcription, adds three atomic pause-analysis profiles, an editor-opening assistant and a cached paragraph-drag preview with an explicit insert target. Version 0.4.3 fixes Update Channel after manually disconnecting generated Movie/Sound pairs by keeping Blender UI polling and document synchronization strictly read-only. Manually separated strips remain associated with their TBE project and may be updated normally.
Newly transcribed projects begin in SOURCE state. Transcribe Channel and transcript import only attach hidden project references: source strips retain their channel, position, length, mute/lock/selection state and native connections. Paragraph split/merge, transcript corrections, pause reanalysis and settings changes also leave the timeline untouched.
The first playback-changing edit—word/filler removal, pause removal or paragraph reorder—atomically creates the non-destructive output. Original strips then become locked, muted masters parked on the highest free channels (normally 127/128), away from the ordinary editing area. Generated Movie/Sound or Sound-only segments on the original channels are the active VSE output. Sequencer preview, Render Animation, command-line rendering and normal exports therefore include text edits before Apply Permanently. A failed first materialization restores the unchanged source timeline and keeps the previous saved project state.
Apply only deletes the protected masters and turns the generated result into ordinary VSE strips. Restore Original Channel removes the generated result and returns the masters. Reset Document Edits restores words, pauses and source paragraph order without deleting transcripts or project binding.
Generated audio cuts use a 10 ms linear crossfade by default. Sound handles may extend a few milliseconds around the logical cut, but the Movie strips, EDL mapping and project duration do not change. A Sound strip first uses its original channel, then lower channels. AV projects may finally use higher free channels, but always strictly below their Movie channel. Audio-only projects use only the original and lower channels. A failed allocation leaves the existing output untouched and identifies the frame and allowed channel range. Advanced → Audio Cut Smoothing enables/disables this behavior and sets a 1–100 ms duration. Apply Permanently retains the resulting volume automation.
Every generated AV segment uses Blender's native Connected Strips feature, so ordinary selection, movement and splitting treat its Movie and Sound together. If no Sequencer area is currently open, the valid rebuilt output is kept and the connection is completed automatically after a Sequencer area becomes available. Manual VSE moves and splits remain in place until the next TBE rebuild; a later transcript edit regenerates the timeline from the TBE EDL. Apply Permanently without an intervening rebuild preserves those manual cuts and the existing native connections.
- In Blender 5.2 choose Edit → Preferences → Get Extensions.
- Open the top-right menu and choose Install from Disk.
- Select
vse_text_based_editing-0.4.3-windows_x64.zip. - Open the VSE Text Based Editing tab in a Video Sequencer sidebar.
The extension requests local file access for media and its persistent cache. Network access is used only when the user downloads a model or worker.
Click a Movie or Sound strip in the Sequencer timeline and press Transcribe Channel.
Transcription creates the document and source/playhead mapping but does not cut the visible channel. The sidebar marks this as “Timeline unchanged until the first removal or reorder.” Update Channel has the same behavior while the project remains in SOURCE state.
- A Movie chooses its unambiguous synchronized Sound and processes every supported Movie on that base channel.
- A Sound with a Movie partner opens that Movie-channel project.
- A Sound without a Movie creates an audio-only Channel Document. Its source may also be an MP4/MOV container with embedded audio.
- Cuts and trims are supported. Reverse, retiming, strip modifiers, overlaps and missing or ambiguous audio partners fail during preflight without changing the timeline.
Later, add another supported strip to the same base channel and use Update Channel. New media is inserted at its timeline position. Changed source files are detected through stored fingerprints and only the affected media transcript is refreshed. Updates are atomic; an error restores the prior output. Complete source transcripts are cached and reused.
One scene can contain independent projects for several base channels. Selecting a generated segment or protected master selects its owning Channel Document.
Use Open Text Based Editor in the Sequencer sidebar to activate the largest available Text Editor (an already active TBE area is preferred). If no Text Editor exists, a short three-step setup explanation is shown; the add-on never changes the layout automatically. You can also change any Blender area to Text Editor and enable Text Based Editor beside the Open control in that area's header. No workspace is created and no existing layout is rearranged, so the Preview, Sequencer and transcript can remain visible together. TBE follows the project belonging to the active Movie or Sound strip. Turning TBE off restores the ordinary Text document that the area showed previously. The area-local TBE mode and Show Removed choice are stored with the Blender screen layout.
Only the explicitly enabled area becomes a controlled transcript document. Other Text Editor areas, normal Text datablocks and Blender's regular Text Editor keymaps continue to work normally. Removed words and pauses are hidden by default, and reorderable sections appear as paragraphs.
- The native blinking caret follows the word at the VSE playhead, and the currently spoken word is highlighted in the Text Editor accent color. Moving the caret or clicking a word seeks the edited timeline.
- Native Text Editor selection can select a contiguous word/pause range.
- While selecting removed text, Blender's native selection color temporarily replaces the gray strikethrough so the complete selection remains visible.
- Delete or Backspace removes a selection.
- Word and pause edits reuse existing generated strips wherever possible, so Preview updates immediately without rebuilding the complete Channel output.
- Removed words use an 80 ms start/end expansion, clamped to their safe word cells. Advanced → Word Cut Accuracy changes the project defaults; Tune Last/Selected Cut stores optional timing overrides for one removed word run. The Whisper timestamps themselves remain unchanged.
- Show Removed in the header reveals removals in gray with a strikethrough; Delete/Backspace on an entirely removed selection restores it.
- Double-click selects the complete word using Blender's native Text Editor behavior. Word correction is not enabled in this release.
- Enter splits before the word at the caret.
- Backspace at the start of a paragraph joins the previous paragraph when no real timeline gap separates them.
- Drag the left
≡handle, or use Alt+Up/Down, to reorder paragraphs across media and gap boundaries. Drag geometry is cached at start; an accent source marker, full-width insert line, pointer label and header text show precisely where the paragraph will land. The project and VSE mutate once, on drop; Escape or dropping outside cancels without changes. - Ctrl+Z / Ctrl+Shift+Z operate document undo/redo.
- Space retains Blender's normal timeline play/pause behavior.
- Word wrapping, scrolling and font-size controls remain the native Blender Text Editor behavior. Free typing, paste and cut are blocked only in TBE mode because the project JSON remains the source of truth.
FFmpeg automatically analyzes each visible source cut during Transcribe Channel
and Update Channel. Acoustic pauses appear as selectable (...) elements
between words, including at paragraph boundaries.
Defaults are 0.30 seconds minimum, -35 dB and 0.05 seconds protection at each
edge. Removing (...) cuts only the interior between those protection edges.
Pauses opens a project menu that can remove enabled detections in one step,
restore every removed pause, and reanalyze the Channel. Minimum duration,
silence threshold and edge protection are editable in the same menu. Raising
the threshold, for example from -35 dB to -30 dB, classifies more and louder
audio as silence. Removed pause ranges remain stored until Apply Permanently,
so Restore Removed Pauses can cut them back in. Real empty spaces between
channel strips are separate, non-editable timeline-gap lines and are never
treated as speech pauses.
Reset Defaults restores the three values from the add-on preferences;
Reanalyze then refreshes the detected pause list with those defaults.
Three one-click profiles set all values and immediately run an atomic analysis:
- Long Gaps: 1.20 s, -38 dB, 120 ms protection per edge
- Between Sentences: 0.60 s, -35 dB, 80 ms protection
- Inside Sentence: 0.30 s, -32 dB, 50 ms protection
Only a successful analysis stores the chosen profile and new detection list; failure or cancellation retains the previous project values and pauses. Manual changes mark the profile as Custom. New detections always start as Present.
Fillers opens the editable project list with Remove Enabled, Restore,
Select Matches and Reset Defaults. Matching is Unicode-aware, case-insensitive
and ignores surrounding punctuation. Defaults include English um, umm,
uh, uhh, uhm, erm, er, ah, hmm, mhm and German äh, ähh,
ähm, ähmm, öh, öhm, hm. Only filler sounds present in the transcript
can be matched.
FFmpeg extracts complete original audio to temporary 16 kHz mono PCM.
whisper-cli.exe runs out of process, so Blender does not load native ML
libraries. Processing is local. Full transcripts live under:
%LOCALAPPDATA%\TextBasedEditing\transcripts\
Cache identity includes normalized source path, size, modification time, model, language, worker version and schema. Models are multilingual Tiny, Base, Small, Medium, Large v3 Turbo and Large v3; Base and English are defaults.
Downloads use .part files, exact catalog sizes and catalogued SHA-256 values,
then install atomically. CDN ETags are ignored. The Model Manager shows
Connecting, Downloading, Verifying and Installing states, transfer progress,
Cancel, persistent success/error feedback and Retry while blocking duplicate
starts.
Schema-6 JSON is stored in hidden .TextBasedEditing.<project-id> Text
datablocks. A scene registry maps independent projects to base channels. Data
includes media fingerprints and caches, media-aware words, paragraph and gap
blocks, full original pause ranges, protection values, the EDL and master/
generated/source strip references plus SOURCE/DERIVED timeline state. Schema-5
projects migrate as DERIVED because they already contain materialized output;
Schema-1/2 single-clip projects and their legacy scene binding migrate
automatically to a one-item Channel Document.
- Windows x64 and Blender 5.2 LTS.
- Forward, unretimed Movie/Sound pairs and audio-only Sound strips.
- No reverse, meta strips, source modifiers, source-strip overlaps, diarization, subtitles, paragraph duplication or free transcript typing.
- Video cuts remain hard; generated audio cuts use configurable linear
crossfades. Apply recovery after saving and closing requires a
.blendbackup.
tests/test_core.py covers schema migration, paragraph/gap behavior, pause
protection, mapping and changed-media replacement. tests/test_runtime.py
covers verified downloads and transcript caches. tests/blender_smoke.py
creates real AV/audio-only channel projects and exercises updates, rollback,
multi-project rendering, Rebuild, Bypass, Restore and Apply.
The extension is GPL-3.0-or-later. It downloads but does not redistribute
whisper.cpp (MIT) or ggml Whisper models. ReScript is a design reference; none
of its application code is embedded.