Skip to content

Latest commit

 

History

History
118 lines (91 loc) · 5.09 KB

File metadata and controls

118 lines (91 loc) · 5.09 KB

Troubleshooting

Common issues and their fixes. If you hit something not listed here, the log file produced by logback (default: stdout, configurable via -Dlogback.configurationFile=...) is usually the right place to start.

"unauthorized" responses from the peer

The peer is rejecting your token.

  • Confirm both daemons see the same token. Each daemon prints its own token on startup if --token was not passed; you can use either side's token, but the UI must enter the peer's token, not its own.
  • Tokens are case-sensitive and 43 characters of URL-safe Base64. Trailing whitespace from a copy-paste will silently break auth.
  • If you restart a daemon without --token, it generates a new random one and the old peer config is no longer valid. Pin both peers to a fixed token in production: --token "$(cat ~/.netcopy/token)".

"connection refused" or "no route to host"

The control plane is on --port (default 7777); the TCP data plane is on --tcp-port (default 7778). Both must be reachable from the peer.

  • Check ss -tlnp | grep -E '(7777|7778)' on the daemon host.
  • Verify the firewall: sudo ufw status (Debian/Ubuntu) or sudo firewall-cmd --list-all (RHEL family). Open the two ports for the peer's source address only — NetCopy assumes a trusted LAN, but a hostile RFC1918 neighbour with the token is still a hostile neighbour.
  • If the daemon is bound to 127.0.0.1, change --bind to the LAN IP or 0.0.0.0.

TCP transfers hang on Windows but HTTP works

Symptom: --tcp-port 7778 accepts the HELLO frame, but data transfers stall indefinitely on a 127.0.0.1 self-test, while --tcp-port 0 (HTTP-only) works fine.

Java NIO Selector on Windows uses an internal self-pipe over loopback. On hosts with aggressive endpoint security (Symantec, CrowdStrike, some WireGuard userspace clients, occasionally Hyper-V virtual switches), this self-pipe wakeup is intercepted in a way that drops or delays the wakeup, which cascades into the connection-handler loop never noticing new data.

Workarounds:

  • Bind to a real LAN IP, not 127.0.0.1.
  • Temporarily disable the AV/VPN to confirm the diagnosis.
  • Run on Linux. Linux is the primary supported platform; Windows is best-effort.
  • Set --tcp-port 0 to fall back to the HTTP data plane, which uses blocking I/O via Javalin/Jetty and does not exercise the same Selector path.

Transfer fails immediately with source_changed

The source file's size or mtime changed between manifest planning and the moment a chunk was pulled. NetCopy refuses to write a half-old, half-new file because the resulting xxh3 hash would no longer match anything meaningful.

  • Re-plan the transfer from the UI; this captures the new size and mtime.
  • If the file is being actively appended to (logs, video recording), finish the recording before transferring, or copy it to a stable location first.

Job is stuck in RUNNING after a kill

You killed the daemon mid-transfer and the next start did not pick the job back up.

  • Check that --state-dir points at the same directory across runs. If it is unset, the default is $XDG_STATE_HOME/netcopy, which on systemd-less setups falls back to ~/.local/state/netcopy — make sure that directory is writable.
  • Look in <state-dir>/jobs/. Each running job should have a <id>.json. If the file exists but the daemon does not log "Resuming job ", check for <id>.json.tmp next to it; that is a half-written state file from a crash mid-fsync. Inspect both, keep the consistent one, and rename it.
  • Sidecars under <receive-root>/.../<file>.netcopy/ are the source of truth for what is already on disk. Even if the job state JSON is lost, re-running the same transfer will pick up where the bitmap left off (provided the source file has not changed).

out of disk space mid-transfer

NetCopy preallocates data.partial to the full target size when a sidecar is created. This is intentional — it surfaces "no space" early instead of five hours into the transfer.

  • Free up space on the receive root and resume the job. The partial file is retained.
  • If you cannot free enough space, cancel the transfer in the UI and manually delete the <file>.netcopy/ directory before re-planning with fewer files.

"port 7777 already in use"

Another process holds the port — most often a stale NetCopy instance.

ss -tlnp | grep :7777
kill <pid>

If you genuinely want two NetCopy instances on the same host, give each its own --port, --tcp-port, and --state-dir.

Browser UI shows nothing or "websocket disconnected"

  • The UI is served from the same --port as the REST API. Hitting http://host:7777/ should return HTML; if it 404s, the JAR was built without the bundled web/ resources (rebuild with ./mvnw -B verify).
  • WebSocket auth uses ?token=... as a fallback. If the browser refuses the connection, open DevTools → Network → WS and confirm the ?token=... query string is present and matches the daemon's token.
  • Some corporate proxies strip WebSocket upgrades. Connect over the LAN directly, or tunnel with ssh -L 7777:localhost:7777 user@deb1.