git-remote-confluence is a Git remote helper that treats a Confluence page
tree or space as a Git remote. It imports Confluence storage-format XML and
attachments into a Git repository and can push committed body updates for
existing pages back to Confluence.
The helper is intentionally narrow: Confluence remains the system that owns page identity, hierarchy, version numbers, and storage-format XML. Git becomes the place where page bodies and sync metadata can be reviewed, edited, and committed.
On fetch or clone, the helper resolves the remote URL to either a page root or a space root, reads pages through the Confluence REST API, and writes a Git fast-import stream. A page root imports that page and its descendants. A space root imports the current pages in that space and reconstructs the hierarchy from ancestor metadata.
On push, the helper scans committed page metadata, finds pages whose stored body content changed, checks the Confluence version and storage hash, and updates the existing Confluence page body. It refuses to overwrite a page when Confluence no longer matches the imported metadata.
Create, delete, move, title, and attachment changes are not pushed yet.
Confluence storage-format XML is the synchronized content format. The companion
git-confluence clean/smudge filter makes that practical for day-to-day
editing:
- Git stores each page body as Confluence storage-format XML.
- The working tree can show the same file as Markdown on checkout.
git addcan convert the edited Markdown back to Confluence storage XML.
That gives users Markdown editing while preserving the storage XML that Confluence needs for reliable import and push.
The imported repository includes this .gitattributes entry:
*.md filter=confluence-storage diff=markdownConfigure the git-confluence filter before checking out imported files.
Each Confluence page is represented by two files:
<pageId>.md
<pageId>.yml
Child pages are placed under their parent's page-id directory:
123456789.md
123456789.yml
123456789/123456790.md
123456789/123456790.yml
Attachments are downloaded below the page ID that owns them:
123456789/attachments/diagram.png
123456789/123456790/attachments/notes.pdf
Path separators and control characters in attachment names are replaced with
underscores so an attachment cannot escape its page's attachments directory.
Confluence grants attachment permissions separately from page permissions, so an
HTTP 403 or 404 on an attachment listing or download never aborts the import,
even on a strict git fetch or a space import. The page is imported without the
unreadable attachments, its metadata records attachments_error.http_status, and
warnings identify the page and summarize the affected count even with --quiet.
HTTP 401/429/5xx and invalid responses still abort the operation.
When a refused attachment was present in the previous import, the warning also names the dropped paths. The import still proceeds as a fast-forward commit, and the earlier commit retains the attachments, so revoked permissions are visible rather than silent.
The .md file is stored in Git as Confluence storage-format XML. With the
git-confluence filter configured, it is checked out as Markdown and converted
back to storage XML on git add.
The .yml file contains page metadata, including the Confluence version number,
links, parent and child page IDs, file paths, and a SHA-256 hash of the stored
XML content for push conflict checks.
Build the helper and put it on PATH:
go build .This writes ./git-remote-confluence. If git --exec-path already contains an
older git-remote-confluence, replace that binary as well because Git may
prefer helpers from its exec path over PATH.
Install the tagged release with Go:
go install github.com/hkwi/git-remote-confluence@v0.1.0Prebuilt archives for Linux, macOS, and Windows are published on the GitHub
Releases page. Each release includes checksums.txt.
Check the installed binary:
git-remote-confluence versionThe helper needs a Confluence personal access token. It reads the first value it finds from:
CONFLUENCE_PATGIT_REMOTE_CONFLUENCE_PATremote.<name>.patconfluence.patremote.confluence.pat
If the git-confluence filter is already configured globally, clone with Git's
explicit remote-helper syntax:
CONFLUENCE_PAT=... git clone \
'confluence::https://confluence.example.com/pages/viewpage.action?pageId=123456789'For a per-clone filter configuration, clone without checkout, configure the filter, then check out:
CONFLUENCE_PAT=... git clone --no-checkout \
'confluence::https://confluence.example.com/pages/viewpage.action?pageId=123456789' \
pages
cd pages
git config filter.confluence-storage.clean "/path/to/git-confluence/git-confluence clean"
git config filter.confluence-storage.smudge "/path/to/git-confluence/git-confluence smudge"
git config filter.confluence-storage.required true
git checkoutThe remote URL may identify a page by pageId, a display page URL, or a
Confluence space.
Initial clones of a page tree allow partial retrieval by default
(confluence.allowPartialClone=true). To require every listed page to be
retrieved, disable partial cloning explicitly:
CONFLUENCE_PAT=... git -c confluence.allowPartialClone=false clone \
'confluence::https://confluence.example.com/pages/viewpage.action?pageId=123456789'With partial cloning enabled, HTTP 403 and 404 responses when fetching a
descendant page's content are skipped, together with that page's subtree.
Other pages and their attachments are still imported. Warnings identify each
skipped page, its parent, and the HTTP status, and summarize the skipped count
even with --quiet. A parent's
metadata lists imported children in children and omitted children in
skipped_children; the skipped count does not include unknown descendants.
A 404 may mean missing content or insufficient permission, so a partial clone
does not establish that the omitted pages have been deleted.
If a readable descendant's child listing returns HTTP 403 or 404, its page and
attachments are retained, along with any children listed in completed batches.
Its metadata records children_error.http_status; children then describes
only the imported subset, not a complete list. Warnings identify the page and
summarize incomplete listings even with --quiet. Unknown descendants are not
reported as a zero-child result.
Root page or root child-list failures, HTTP 401/429/5xx, invalid responses, and persistent network failures still abort the operation. Space imports also retain their existing strict behavior.
The option accepts true or false and can also be set via, in precedence
order, CONFLUENCE_ALLOW_PARTIAL_CLONE,
GIT_REMOTE_CONFLUENCE_ALLOW_PARTIAL_CLONE, remote.<name>.allowPartialClone,
confluence.allowPartialClone, or remote.confluence.allowPartialClone.
It applies only when Git identifies the operation as cloning. Subsequent
git fetch operations remain strict even if the setting persists, so a
retrieval failure cannot replace an existing snapshot with a partial one.
Fetching an already-cloned Confluence mapping chains each import onto the
existing local tip with a from line, so a successful fetch is a fast-forward
update rather than a rewrite. This lets git fetch refresh a snapshot repeatedly
without being rejected as non-fast-forward.
GET requests retry temporary connection failures, including refused proxy connections, resets, and timeouts before a response is received. There are at most four attempts, with waits of 1, 2, and 4 seconds. Each attempt retains the 60-second HTTP timeout. Progress output shows the failed request, wait, and next attempt; only that request is retried, so completed page requests are not restarted.
HTTP error responses, invalid JSON, and failures while reading response bodies
are not retried. PUT requests are not automatically resent. If retries are
exhausted, the import fails and existing Git refs remain unchanged, including
when partial cloning is enabled. Git's stream ends early message then reflects
the incomplete import; the preceding helper error identifies the cause.
By default, REST requests use the traditional unversioned /rest/api root. If
a Confluence installation exposes the API below a different root or requires an
explicit version, set either or both of these variables when cloning:
CONFLUENCE_API_ROOT=custom/api/root \
CONFLUENCE_API_VERSION=2.0 \
CONFLUENCE_PAT=... \
git clone 'confluence::https://confluence.example.com/pages/viewpage.action?pageId=123456789'This example sends content requests below
/custom/api/root/2.0/content/.... CONFLUENCE_API_VERSION is omitted by
default, preserving /rest/api/content/.... The aliases
GIT_REMOTE_CONFLUENCE_API_ROOT and GIT_REMOTE_CONFLUENCE_API_VERSION are
also accepted.
For an existing configured remote, use Git configuration instead:
[remote "origin"]
apiRoot = custom/api/root
apiVersion = 2.0The fallback keys are confluence.apiRoot, confluence.apiVersion,
remote.confluence.apiRoot, and remote.confluence.apiVersion. Download and
browser links returned by Confluence remain relative to the site URL; the API
root and version are not added to them.
After editing and committing page Markdown, push existing page body updates back to Confluence:
git push origin HEAD:main
git fetch originFetch after a successful push to refresh the Confluence page version and local metadata.
For a configured remote, confluence:https://... is accepted by the helper when
Git is told to use the confluence VCS helper:
[remote "origin"]
vcs = confluence
url = confluence:https://confluence.example.com/pages/viewpage.action?pageId=123456789
pat = somevalueTo show helper progress, ask Git for progress or verbose output:
CONFLUENCE_PAT=... git clone --progress --verbose \
'confluence::https://confluence.example.com/pages/viewpage.action?pageId=123456789'Progress is written to stderr. When Git captures helper stderr during a successful import, the helper also mirrors progress to the controlling terminal if one is available.
go test ./...