Skip to content

neuromorphs/AudioVisual_Engagement

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

17 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

AudioVisual_Engagement

A cross-modal (audio vs. visual) selective-attention paradigm in PsychoPy, designed for EEG. A subject chooses which stream to attend, watches a muted gameplay video while a continuous speech clip plays, and then answers a yes/no attention-check about the attended stream.

The stimulus library is prepared once and reused every session — no audio/alignment is regenerated at run time.

Layout

File Purpose
config.yaml All parameters (paths, timing, dataset, probes). Edit this, not the code.
build_stimuli.py Run once. Downloads real speech + word onsets into stims/ and writes stims/manifest.json.
run.py The launcher — collects subject/session metadata and runs a session into data/<subject>/<session>/.
run_block.py The session engine (called by run.py; also runnable directly for quick tests / --test-trigger).
reconstruct_tetris.py Renders the saved Tetris games to MP4 after a session.
data/ Per-subject / per-session output (see below). (git-ignored)
cma_common.py Shared config / manifest / path helpers.
cma_webcam.py Separate-process webcam recorder (used by run_block.py).
tetris/ Self-playing black-and-white Tetris (the default visual stimulus).
build_video_segments.py (Only for visual.mode: video) detects the game segments of the recorded clip.
stims/ Prepared audio (*.wav + *.words.json/csv) + manifest.json. (git-ignored)
behavior/ Per-subject behavioural logs (<subject>.csv + per-block JSON). (git-ignored)
webcam/ Per-block webcam recordings + frame-timestamp sidecars. (git-ignored)
stim_files/tetris.mp4 The visual stimulus.

Install

source .venv/bin/activate            # Python 3.11
python -m pip install -r requirements.txt
# (or pin exact versions with requirements.lock.txt)

1. Build the stimulus library (once)

python build_stimuli.py              # builds build.n_clips (default 10) clips
python build_stimuli.py --add 5      # append 5 more later
python build_stimuli.py --list       # show what's prepared
python build_stimuli.py --source tts # alternative: local gTTS + faster-whisper

Audio source: gilkeyio/librispeech-alignments — real LibriSpeech audiobook speech with Montreal-Forced-Aligner word onsets (CC-BY-4.0). Same-speaker utterances are concatenated to ~30 s clips; word onsets are time-shifted accordingly. Each clip gets an auto-generated yes/no word probe (one true-target and one false-target variant).

1b. Detect the usable video segments (once, video mode)

python build_video_segments.py            # writes visual.video.segments_file
python build_video_segments.py --preview  # also print each segment

The tetris clip has non-game (colour) sections that must not be shown. Gameplay is black-and-white, so this classifies frames by colour saturation and writes the contiguous grayscale intervals long enough to hold a block. Random cuts are then drawn ONLY from those segments (currently ~74 % of the clip; colour menus/intros are skipped). Re-run it if you change the video or visual.video.filter.

2. Run a session

Use the launcher — it collects participant/session metadata (a GUI dialog, or --no-gui + flags) and organises everything under data/<subject>/<session>/:

python run.py                                    # GUI dialog for the metadata
python run.py --subject P01 --session 01         # prefilled dialog
python run.py --subject P01 --session 01 --no-gui \
              --age 24 --sex female --handedness right --experimenter MA
python run.py --subject P01 --session 01 --trials 4 --no-gui   # short test

Each session is written to (and will not overwrite an existing one without --force):

data/<subject>/<session>/
    session_info.json     participant + session metadata + provenance (git commit,
                          seed, config file, software versions)
    behavior/             <subject>.csv + per-session JSON (all trial data)
    games/                Tetris reconstruction records  ->  reconstruct_tetris.py
    webcam/               webcam recording(s)

run_block.py is the underlying engine and can still be run directly for quick tests (writes to the flat behavior/, games/, webcam/ defaults) and for --test-trigger.

One session runs experiment.n_trials (default 60) trials back-to-back. Per trial: focus cue/choicejittered fixation delay (base gap_seconds, default 2 s, + Gaussian jitter) → synchronous visual stimulus + speech clip (PTB flip-scheduled onset, experiment.block_seconds long, default 20 s) → attention probe about the attended stream → an inter-trial fixation. Audio clips are assigned in a balanced shuffle. Results: one row per trial in behavior/<subject>.csv plus a per-session JSON.

Timing jitter (experiment.jitter): the pre-stimulus delay is jittered per trial to decorrelate onsets for EEG — Gaussian std_ms (default 150), clamped to ±max_ms (default 300). The log records the actual delay_seconds, delay_jitter_ms, delay_frames, and the instruction_seconds (measured length of the cue/choice screen).

Focus assignment (experiment.focus_mode):

  • auto (default) — the experiment assigns focus, balanced 50/50 and shuffled (30 Audio / 30 Visual for 60 trials), and cues the subject ("Attend to the AUDIO") for cue_seconds.
  • manual — the subject chooses each trial (← Audio / → Visual).

Photodiode trigger (trigger): a white square in the bottom-right corner is shown for the whole 1 s gap — its rising edge = gap onset, falling edge = audiovisual onset — for hardware-precise EEG timing. Size/corner configurable.

Visual stimulus (visual.mode)

  • tetris (default) — a self-playing, black-and-white Tetris (in tetris/), watched passively. Deterministic (seeded), grayscale (no colour confound), compact and centred with a fixation dot to limit eye movements. The attend-visual task is to report the greatest height the stack reached (tetris_max_height), scored automatically. Speed is visual.tetris.speed_s_per_step. (Preview: python tetris/tetris_game.py --save-frames 6 writes frames to /tmp.) Every Tetris trial is saved to games/ as a reconstruction record (seed + params + block duration + exact display frame times); because the game is deterministic these render to faithful MP4s after the session:

    python reconstruct_tetris.py                 # all trials -> games/mp4/*.mp4
    python reconstruct_tetris.py --exact         # frame-exact (uses recorded times)

    Reproducibility: the master seed is fixed in experiment.seed (constant across runs), and the full config is embedded in each session JSON.

  • gabor — a small, centred Gabor patch that rotates and randomly reverses direction; attend-visual counts the reversals (numeric probe, auto-scored).

  • video — the pre-recorded tetris.mp4 with random cuts from the black-and-white game segments (needs build_video_segments.py; attend-visual trials are logged unscored — see the caveat below).

All three keep the eyes centred; the auto cue reminds the subject of the task ("Count the row clears" / "reversals" / "Listen for the words").

The auditory task (either mode) is a yes/no "was WORD spoken?" probe whose target is chosen at trial time from words actually heard within the played window — so it stays valid even though only the first block_seconds of each ~30 s clip is played.

Webcam (optional)

If webcam.enabled is true in config.yaml, each block records the participant to webcam/<subject>_<ts>_b<block>.mp4 in a separate process (so it never disturbs the draw-loop timing). Alongside the video it writes: <name>.frames.csv — a wall-clock timestamp per frame — and <name>.status.json. The behaviour log stores webcam_av_onset_wallclock, so you can find the exact frame at audiovisual onset for EEG alignment.

macOS: grant Camera access to your terminal/Python under System Settings → Privacy & Security → Camera (you'll be prompted on first run). If the camera can't open and webcam.required is false, the block runs without recording.

⚠️ Visual attention-check in video mode with random cuts

With visual.video.random_start: true, every trial shows a different random cut, so a single pre-filled answer in visual.video.probes[].correct can't be correct for all of them. As shipped, attend-visual trials are logged UNSCORED (probe_correct = null) in this mode — attend-audio trials are scored normally. Options if you need a scored visual check:

  • switch to visual.mode: gabor (reversal count, auto-scored, also fixes the eye-movement confound); or
  • turn off random cuts (random_start: false) and fill probes[].correct for the fixed cut; or
  • ask for a per-cut visual probe (needs a way to know the ground truth per cut).

About

No description, website, or topics provided.

Resources

License

Stars

1 star

Watchers

0 watching

Forks

Releases

No releases published

Packages

 
 
 

Contributors

Languages