SpringTorch is an application composed of a Spring Boot backend and a Python worker for running Qwen inference through Temporal.io.
Spring Boot exposes the HTTP API and orchestrates Workflow execution. The Python worker does not expose a web endpoint: it listens to a Temporal Task Queue and runs the inference Activity using PyTorch.
HTTP Client
|
v
Spring Boot :8080
|
v
Temporal :7233
|
v
Task Queue: llm-task-queue
|
v
Python Worker + Qwen
Services defined in docker-compose.yml:
backend: Spring Boot API and Temporal Java Worker.jobs: Python Worker that loads Qwen and runs inference.postgres: PostgreSQL database for Spring.temporal: Temporal development server with Web UI.
The backend is located in backend/ and uses:
- Spring Boot
4.1.1. - Kotlin
2.4.20. - JDK
21. - Gradle Kotlin DSL.
- Spring Data JPA.
- Flyway.
- PostgreSQL Driver.
- SpringDoc OpenAPI.
- HTMX Spring Boot.
- Testcontainers for PostgreSQL integration tests.
- Temporal Java SDK
1.38.0.
The backend contains the Temporal Workflow, the Temporal client, the Java Worker that executes the Workflow, and the REST controller.
The worker is located in jobs/ and uses:
- Python
3.14. uvfor dependency resolution and installation.- PyTorch.
- Transformers.
- Accelerate.
- Temporal Python SDK (
temporalio). Qwen/Qwen2.5-0.5B-Instructby default.
The worker listens to the llm-task-queue Task Queue. It does not use FastAPI, Flask, or any other HTTP framework.
The image installs build-essential because Triton may need to compile code during inference. The model is stored in the qwen-model-cache Docker volume to avoid downloading it every time the container is recreated.
You can select another model using QWEN_MODEL:
docker compose run --rm -e QWEN_MODEL=Qwen/Qwen2.5-1.5B-Instruct jobsTemporal runs in development mode through the temporal service in Compose:
- Server:
localhost:7233. - Web UI:
http://localhost:8233. - Persistence: in-memory, suitable for local development.
Within the Docker network, services connect using:
temporal:7233
The backend and worker receive this address through TEMPORAL_ADDRESS.
The text generation endpoint is:
POST http://localhost:8080/api/llm/generate
Content-Type: application/jsonRequest body:
{
"prompt": "Explain what Temporal.io is in a few words"
}The response has the following shape:
{
"text": "..."
}The complete flow is:
- Spring receives the HTTP request.
- Spring starts a Temporal Workflow.
- The Workflow schedules the
generate_qwen_responseActivity onllm-task-queue. - The Python worker receives the Activity.
- Qwen generates the response.
- Temporal returns the result to the Workflow, and Spring responds to the client.
Swagger UI:
http://localhost:8080/swagger-ui/index.html
OpenAPI specification:
http://localhost:8080/v3/api-docs
Build and start all services:
docker compose up --buildRun in the background:
docker compose up -d --buildStop the services:
docker compose downTo also remove PostgreSQL data and the persistent Qwen cache:
docker compose down -vThe jobs service requests an NVIDIA GPU through Compose. To run inference on the GPU, Docker Desktop must have the corresponding NVIDIA support available.
Each component has its own Dockerfile:
backend/Dockerfile: multi-stage build with Gradle/JDK 21 and an Eclipse Temurin JRE 21 runtime.jobs/Dockerfile:uvimage with Python 3.14, worker dependencies, and the build support required by Triton.
The configuration is in .devcontainer/devcontainer.json and uses docker-compose.yml with backend as the primary service:
- Workspace mounted at
/workspaces/SpringTorch. - Port 8080 forwarded.
- Remote user
root. - PostgreSQL and Temporal available as Compose services.
docker-outside-of-dockerfeature enabled.
Docker-outside-of-Docker allows Docker commands to run from the Dev Container through the Docker socket. This is required for Testcontainers to create and manage containers during tests.
To open the environment:
- Install the VS Code Dev Containers extension.
- Open the project folder.
- Run
Dev Containers: Rebuild and Reopen in Container.
Build the backend from backend/:
cd backend
.\gradlew.bat buildRun the tests:
cd backend
.\gradlew.bat testUpdate Python dependencies:
cd jobs
uv lock
uv syncThe jobs/uv.lock file keeps the resolved versions for the Python worker.
- Docker Desktop with Docker Compose.
- NVIDIA support in Docker Desktop to run Qwen on the GPU.
- VS Code and the Dev Containers extension if using the development environment.
- JDK 21 only if building the backend outside Docker.
uvand Python 3.14 only if running the worker outside Docker.