Control Windows virtual desktops with hand gestures, using the webcam you already have. It runs as a real tray application — no browser needed, no cloud, no account — and drives the operating system's own keyboard shortcuts rather than simulating a desktop manager of its own.
Swipe Left → Previous desktop Win + Ctrl + Left
Swipe Right → Next desktop Win + Ctrl + Right
Open Palm → Task View Win + Tab
Fist → Minimise window Win + Down
Two Fingers → New desktop Win + Ctrl + D
The hard part of a gesture controller is not recognising gestures. It is not recognising them — not switching your desktop because you reached for a mug, waved at a colleague, or talked with your hands. Most of this codebase is about that.
| OS | Windows 10 (1903+) or Windows 11 — virtual-desktop shortcuts are Windows-only |
| Python | 3.10 – 3.12 (MediaPipe has no Windows wheel for 3.13 yet) |
| Camera | Any webcam. 640×480 at 30 fps is plenty |
| GPU | Not required. The landmark model runs comfortably on CPU |
| Network | Not required, ever |
git clone https://github.com/PremPastagia/Gesture-Control.git
cd Gesture-Control
install.batinstall.bat creates venv\, installs the dependencies, and downloads the
MediaPipe hand-landmark model if it is missing. To do it by hand:
python -m venv venv
venv\Scripts\python -m pip install -r requirements.txtrun.bat REM start in the tray
venv\Scripts\python src\main.py REM same thing, with a console for errors
venv\Scripts\python src\main.py --minimizedThe first launch opens a short setup: it checks the camera, teaches you the five gestures, and lets you practise them in a mode where nothing is executed. Gesture recognition stays off until you switch it on. The app never starts controlling desktops on its own.
venv\Scripts\python src\build.py --cleanOutput lands in dist\GestureControl\. It is a one-folder build on purpose: a
one-file bundle would unpack ~200 MB of MediaPipe and OpenCV on every launch, which
is the wrong trade for an app that starts at login. settings.json and logs\ are
written next to the executable, so keep that folder writable.
The icon colours tell you the state at a glance — green on, grey off, amber paused, blue test mode, red camera problem — and the first item in its menu toggles recognition. You never have to open a window to switch gestures off. The right-click menu also offers Settings, Calibration, Test-without-executing, timed pauses, and Start with Windows.
Ctrl + Alt + G pauses and resumes from anywhere, even with every window closed.
It is configurable in Settings.
A swipe has to look deliberate:
- Hold your hand still for about half a second. This is the single most important filter — a hand already in motion cannot start a swipe, which is what stops waving from switching desktops.
- Sweep sideways about one and a half palm widths, decisively, in one direction.
- That's it. The action fires, then a cooldown starts and your hand must settle again before the next gesture counts.
A pose (palm, fist, two fingers) has to be held still for several consecutive frames. A pose glimpsed mid-wave never fires.
If you are not sure why something did or did not register, open Calibration. It shows every measurement against the threshold it has to clear, and names the exact reason the last gesture was rejected.
One slider, 0–100, default 60. It moves everything together: swipe distance, swipe speed, confidence, horizontal-versus-vertical ratio, direction consistency, pose hold time, cooldown, and how settled the hand must be before a swipe. Higher reacts to smaller movements and misfires more often; lower demands deliberate gestures and almost never does.
Measured on a synthetic wave-versus-swipe grid (tests/test_state_machine.py):
| Sensitivity | False triggers from waving | Deliberate swipes detected |
|---|---|---|
| 40 | 0 | 20 / 25 |
| 60 (default) | 0 | 25 / 25 |
| 80 | 28 | 25 / 25 |
| 100 | 84 | 25 / 25 |
That table is the whole design argument for the default: at 60 the recogniser catches every deliberate swipe in the grid and nothing from four seconds of continuous waving at any frequency or amplitude tested.
If you want exact numbers instead, Settings → Advanced thresholds replaces the slider with the individual values.
Layers, in the order a gesture meets them:
| Layer | What it rejects |
|---|---|
| Hand scale bounds | A hand too close to the lens or too far to track reliably |
| Full visibility | A hand half outside the frame |
| One hand only | Two hands in view — never read as a repeat of one |
| Tracking confidence | Frames the landmark model is unsure about |
| Pose ambiguity | A hand shape that fits two templates almost equally well |
| Launch from rest | A swipe that began while the hand was already moving |
| Distance | Movement shorter than the threshold, measured in palm widths |
| Velocity | Slow repositioning |
| Axis ratio | Movement that drifted vertically as much as horizontally |
| Direction consistency | A trajectory that changed direction part way |
| Confidence | A match below the required score for that gesture |
| Hold frames | A pose that did not survive N consecutive frames |
| Steadiness | A pose held while the hand was travelling |
| Rate limit | More actions per second than allowed, whatever was recognised |
| Cooldown | Anything during the cooldown after an action |
| Neutral state | Anything before the hand settles and leaves the previous pose |
| Safe Mode | Desktop-changing gestures held to strictly higher thresholds |
| Context gate | Blocklisted app focused, or a borderless full-screen app running |
| Global enable / pause | The tray toggle and the hotkey, re-checked at dispatch |
Two of these deserve a note.
Distances are in palm widths, not pixels. Every spatial threshold is normalised by the distance between the index and pinky knuckles. The same physical hand movement counts the same whether you are leaning into the laptop or sitting back, which is what makes one sensitivity setting work across seating positions.
One gesture is one action. After an action the trajectory history is cleared, a cooldown runs, and the hand must return to a neutral state before anything else can fire. A single long sweep across the whole frame produces exactly one desktop switch, never a chain of them.
camera.py grabber thread, one frame in flight, mirrored + RGB at source
↓
hand_tracker.py MediaPipe Hand Landmarker, LIVE_STREAM mode, async submit
↓
gestures.py rotation/scale-invariant finger scoring, binary encoding,
static gesture classification with a real confidence
↓
state_machine.py IDLE → HAND_DETECTED → GESTURE_CANDIDATE → GESTURE_CONFIRMED
→ ACTION_EXECUTED → COOLDOWN → WAIT_FOR_NEUTRAL → IDLE
↓
actions.py allowlisted Windows shortcuts, one batched SendInput call
↓
tray.py / main.py tray icon, global hotkey, settings window, lifecycle
Supporting modules: config.py (nested schema, atomic writes, migration),
windows_context.py (foreground app, full-screen detection, virtual-desktop
awareness, run-at-login), pipeline.py (sequencing and measurement),
ui_server.py (local REST/WebSocket API), feedback.py, logger.py, perf.py.
Thread layout: the settings window owns the main thread (pywebview requires it on Windows), while the tray, the hotkey message loop, the camera grabber, the recognition pipeline and the HTTP server each run on their own. Nothing in the gesture path waits on the UI.
Measured with tests/benchmark.py:
landmark geometry + classification mean 0.0122 ms p95 0.0130 ms
gesture state machine mean 0.0206 ms p95 0.0427 ms
action dispatch mean 0.0023 ms p95 0.0023 ms
-------------------------------------------------------------------
confirmed gesture → shortcut dispatched (p95 sum) 0.058 ms
So the code in this repository adds well under a millisecond between "the gesture is confirmed" and "the keystroke is handed to Windows".
The honest caveat, which the plan asked for explicitly: the full camera-to-action path cannot be under 30 ms and this project does not claim it is. Webcam exposure and USB transfer alone are typically 15–35 ms, and MediaPipe inference adds a few more. What is optimised here is everything after the frame arrives:
- one frame in flight, never a queue — stale frames are dropped, not processed
CAP_PROP_BUFFERSIZE=1and DirectShow, which measure faster than Media Foundation on most consumer webcams- colour conversion and mirroring done once, at capture
- asynchronous inference — submitting a frame never blocks
- the whole key chord sent in a single
SendInputcall, with no sleep between press and release - frame rate throttled to 12 fps when no hand is in view
Run python tests/benchmark.py --live on your own machine for the real end-to-end
figures, and read them on the Performance tab while the app is running.
venv\Scripts\python -m unittest discover -s tests
venv\Scripts\python tests\benchmark.py --live77 tests, and they run on any OS — no camera, no MediaPipe, no Windows required. Hands
are synthesised by a small forward-kinematic model (tests/hand_fixtures.py) and fed
through the real geometry and state-machine code, and the pipeline tests swap in a fake
camera, a fake tracker and a recording key sender so the whole chain from frame to
Windows shortcut can be asserted on.
What they cover: valid and invalid swipes, waving at eight frequencies, jittery drift, diagonal movement, slow movement, low-confidence and disappearing hands, two hands, hand too close and too far, partial visibility, ambiguous poses, one-gesture-one-action, cooldown, neutral state, rate limiting, per-gesture overrides, Safe Mode strictness, the blocklist, full-screen pause, test mode, camera release when switched off, config migration and clamping, and rejection of unknown actions.
- Every frame is processed in this process on this computer.
- No frame is uploaded, and none is written to disk under any setting.
- The camera device is released — not merely ignored — whenever recognition is off and nothing is previewing.
- Diagnostics exports contain measurements, settings and event names only. The application blocklist is stripped from them, since it can name software you would rather not disclose in a bug report.
- Nothing is registered outside
HKCU\...\Run, and only if you enable Start with Windows. Administrator rights are never requested.
Gestures can only trigger the eight shortcuts in ACTION_SHORTCUTS. That list is the
complete set of things this application can do to your computer. A hand-edited
settings.json naming anything else is rejected at load, not at execution.
TROUBLESHOOTING.md— camera problems, hotkey conflicts, gestures not registering, packaging issuesPlan.txt— the original specification this was built against
- Hand landmarks: MediaPipe Hand Landmarker (Google, Apache 2.0)
- The binary finger-encoding scheme for static poses follows Vidhyanjali et al., Hand Gesture Control Apps Using Hand Movements, Journal of Neonatal Surgery 14(32s), 2025 — included in this repository.
- Landmark normalisation relative to the wrist and palm scale follows the approach in kinivi/hand-gesture-recognition-mediapipe.