A full-stack LLM comparison and benchmarking workspace built with Spring Boot, Spring AI, and React.
Broadcast a single prompt to multiple frontier and local AI models concurrently, evaluate side-by-side outputs, and measure real-time latency and first-response performance.
- π Live Frontend Demo: https://spring-ai-studio-psi.vercel.app/
- π¦ GitHub Repository: https://github.com/Mohammad-Asfin/Spring-AI-Studio
- π― Project Overview
- β¨ Key Features
- π§ Why Spring AI?
- ποΈ System Architecture
- π Request / Response Flow
- β‘ Parallel LLM Execution
- π First Response Detection
- π LLM Benchmarking & Latency Tracking
- π Project Structure
- β Backend Architecture
- π¨ Frontend Architecture
- π οΈ Technology Stack
- π API Documentation
- π Environment Variables
- π» Local Development Guide
- π€ AI Provider Setup
- π Backend Deployment Guide
- βοΈ Vercel Frontend Deployment
- π End-to-End Production Wiring
- π Security & Best Practices
- π§ͺ Testing & Verification
- π Troubleshooting Guide
- π§© Common Use Cases & Prompts
- π Performance Considerations
- πΈ Screenshots
- πΊοΈ Future Roadmap
- π€ Contributing
- π License
- π Important Documentation Links
Spring AI Studio is a full-stack, developer-oriented LLM evaluation workspace and benchmarking suite. It allows users to enter one single prompt and simultaneously query multiple frontier cloud models and local self-hosted open-weight models:
- OpenAI (
GPT-4ocloud model) - Anthropic (
Claudecloud model) - Ollama / DeepSeek (
deepseek-r1:14blocal inference engine)
When selecting Large Language Models for applications, engineers face critical architectural and operational trade-offs:
- Cloud vs. Local Models: Commercial cloud APIs offer cutting-edge reasoning but introduce network latency and recurring token costs. In contrast, local models running via Ollama provide zero API fees and complete data privacy, but depend on local hardware compute.
- Provider Fragmentation: Each AI provider traditionally requires proprietary client SDKs, incompatible JSON formats, and disparate error schemes.
- Subjective vs. Empirical Evaluation: Testing prompts manually across multiple web tabs is tedious and subjective. Spring AI Studio provides an empirical side-by-side view with synchronized request dispatch and high-resolution latency measurement.
React 19 (Vite)
β
β HTTP POST (JSON: { "prompt": "..." })
βΌ
Spring Boot 3.4 REST Controllers
β
βΌ
Spring AI ChatClient Unified Abstraction
β
βββββ΄ββββββββββββββββ¬βββββββββββββββββββββββ
βΌ βΌ βΌ
OpenAI API Anthropic API Local Ollama Daemon
(GPT-4o) (Claude) (deepseek-r1:14b)
β β β
βββββββββββββββββββββΌβββββββββββββββββββββββ
βΌ
Asynchronous Text Responses
β
βΌ
Real-Time Benchmark UI & Latency Telemetry
- Side-by-Side Multi-LLM Comparison: Broadcast a single prompt across OpenAI GPT-4o, Anthropic Claude, and local Ollama DeepSeek concurrently.
-
Unified Spring AI Backend: Clean Spring Boot 3.4 architecture utilizing the official
spring-ai-bom(1.0.0-M6) with zero provider-specific boilerplate in controllers. - Non-Blocking Parallel Client Dispatch: React fires concurrent HTTP requests; slow, queuing, or failing models never block responses from faster providers.
-
Individual Model Lifecycle States: Every model card independently transitions through
IDLE,LOADING(with provider-themed pulsing animations),SUCCESS, andERRORstates. -
Race-Condition-Safe First Response Winner: Uses React
useRefto reliably capture the first successful response (β‘ First Response) across concurrent execution streams. -
High-Resolution Response Timing: Browser
performance.now()measures elapsed execution latency to two decimal places in seconds. -
Live Aggregated Benchmark Statistics: Instant header dashboard displaying:
-
Models Tested (
$N=3$ ) - Successful Responses
- Failed Responses
- First Response Provider
- Average Response Time (calculated across successful runs)
-
Models Tested (
-
Actionable Diagnostic Error Handling: Missing API keys or an offline Ollama daemon render helpful diagnostic badges (e.g.
Set SPRING_AI_OPENAI_API_KEYorRun on http://localhost:11434) instead of crashing. - One-Click Markdown/Text Copying: Quick clipboard export on all generated model cards with visual confirmation.
- Curated Prompt Preset Library: Pre-configured chips for instant testing of technical concepts (Dependency Injection, REST vs GraphQL, Java Streams, Spring AI).
-
Dual-Theme Engine: Native Light and Dark modes built with CSS custom properties and persisted in
localStorage. - Decoupled Deployment Architecture: React single-page frontend ready for Vercel edge CDN and Spring Boot backend ready for JVM container platforms (Railway, Render, AWS).
Prior to Spring AI, integrating multiple LLM providers in Java applications required adding disparate third-party libraries, managing inconsistent API client specifications, and manually transforming request/response POJOs for each vendor.
Spring AI introduces a standardized abstraction over artificial intelligence models, similar to how Spring Data abstracts relational and NoSQL databases. The core abstraction is the ChatClient and ChatModel interface.
// Common unified pattern across OpenAI, Anthropic, and Ollama
String response = chatClient.prompt(userPrompt)
.call()
.content();Whether connecting to OpenAI, Anthropic, or an on-premise Ollama instance, the business logic remains uniform. Switching or adding new AI models (e.g. Mistral, Google Gemini, Amazon Bedrock) requires minimal configuration changes rather than refactoring the codebase.
Spring AI starter dependencies (spring-ai-openai-spring-boot-starter, spring-ai-anthropic-spring-boot-starter, spring-ai-ollama-spring-boot-starter) automatically wire API credentials, base URLs, and HTTP client pooling via standard Spring Boot application.properties.
flowchart TD
subgraph Client["Frontend Client (React 19 + Vite 6.2)"]
UI["Prompt Workspace & Interactive UI"]
Theme["Theme Engine (Light / Dark)"]
Tracker["Latency & First-Response Telemetry Engine"]
end
subgraph Server["Spring Boot 3.4.3 REST Backend (Port 8080)"]
OpenAICtrl["OpenAIController\nPOST /api/openai/ask"]
AnthropicCtrl["AnthropicController\nPOST /api/anthropic/ask"]
OllamaCtrl["OllamaController\nPOST /api/ollama/ask"]
ChatClientLayer["Spring AI ChatClient Layer"]
end
subgraph External["AI Providers & Local Daemons"]
OpenAISvc["OpenAI API\n(GPT-4o)"]
AnthropicSvc["Anthropic API\n(Claude)"]
OllamaDaemon["Local Ollama Daemon\n(deepseek-r1:14b on :11434)"]
end
UI -->|"Parallel HTTP POST { prompt }"| OpenAICtrl
UI -->|"Parallel HTTP POST { prompt }"| AnthropicCtrl
UI -->|"Parallel HTTP POST { prompt }"| OllamaCtrl
OpenAICtrl --> ChatClientLayer
AnthropicCtrl --> ChatClientLayer
OllamaCtrl --> ChatClientLayer
ChatClientLayer -->|"HTTPS API Call"| OpenAISvc
ChatClientLayer -->|"HTTPS API Call"| AnthropicSvc
ChatClientLayer -->|"HTTP localhost:11434"| OllamaDaemon
OpenAISvc -.->|"Generated Text"| OpenAICtrl
AnthropicSvc -.->|"Generated Text"| AnthropicCtrl
OllamaDaemon -.->|"Generated Text"| OllamaCtrl
OpenAICtrl -.->|"Text + Latency"| Tracker
AnthropicCtrl -.->|"Text + Latency"| Tracker
OllamaCtrl -.->|"Text + Latency"| Tracker
Tracker --> UI
- Presentation Layer (React 19 + Vite 6.2): Handles user interaction, theme switching, prompt validation, concurrent network orchestration, and real-time benchmark calculations.
- API & Routing Layer (Spring Boot Web): Exposes stateless REST endpoints mapped under
/api/*, accepts JSON payloads, and handles HTTP status codes. - AI Orchestration Layer (Spring AI 1.0.0-M6): Uses
ChatClient.create(chatModel)to inject provider-specific drivers (OpenAiChatModel,AnthropicChatModel,OllamaChatModel) into unified execution pipelines. - Inference Layer:
- OpenAI / Anthropic: External cloud infrastructure over secure HTTPS.
- Ollama: Self-hosted local daemon running on
http://localhost:11434.
sequenceDiagram
autonumber
actor User
participant React as React 19 Frontend
participant Boot as Spring Boot 3.4
participant SpringAI as Spring AI ChatClient
participant Providers as AI Providers (OpenAI, Claude, Ollama)
User->>React: Enters prompt and clicks "Compare Models"
Note over React: 1. Set all cards to LOADING<br/>2. Start performance.now() timers<br/>3. Reset firstModelRef = null
par Parallel HTTP Requests
React->>Boot: POST /api/openai/ask { "prompt": "..." }
React->>Boot: POST /api/anthropic/ask { "prompt": "..." }
React->>Boot: POST /api/ollama/ask { "prompt": "..." }
end
Boot->>SpringAI: chatClient.prompt(prompt).call()
SpringAI->>Providers: Execute provider-specific HTTP call
Providers-->>SpringAI: Return raw model text response
SpringAI-->>Boot: Extract content string / ChatResponse
par Independent Asynchronous Responses
Boot-->>React: 200 OK (OpenAI Response Body)
Note over React: 1. Calculate OpenAI elapsed time<br/>2. Check firstModelRef: Lock OpenAI as β‘ First Response<br/>3. Transition OpenAI card to SUCCESS
and
Boot-->>React: 200 OK (Claude Response Body)
Note over React: 1. Calculate Claude elapsed time<br/>2. firstModelRef already locked (skip)<br/>3. Transition Claude card to SUCCESS
and
Boot-->>React: 200 OK (Ollama Response Body)
Note over React: 1. Calculate Ollama elapsed time<br/>2. Transition Ollama card to SUCCESS
end
Note over React: Compute aggregate benchmark metrics (Tested: 3, Success: 3, Failed: 0, Avg Time: ~1.45s)
React->>User: Render side-by-side cards with latency metrics and gold winner badge
- User Input: User enters a prompt or clicks an example chip.
- Client Validation: Frontend ensures prompt is non-empty.
- State Initialization: Frontend sets all 3 model cards to
LOADING, starts timers, and resets the first-responder tracker. - Parallel Dispatch: The browser fires three concurrent asynchronous
fetch()requests to Spring Boot. - Controller Processing: Backend
@RestControllervalidates the JSON payload{"prompt": "..."}. - ChatClient Call: Spring AI
ChatClientdelegates the prompt to the respectiveChatModel. - Inference Execution: Cloud APIs or local Ollama process token generation.
- Backend Return: Spring Boot returns
200 OKwith the plain text response, or500on error. - Latency Capture: Frontend calculates elapsed time via
((performance.now() - startTime) / 1000).toFixed(2). - Card Transition: Individual model card transitions from
LOADINGtoSUCCESSorERROR. - First-Response Lock: The first model to resolve successfully locks
firstModelRefand receives the winner badge. - Benchmark Aggregation: Header stats bar recalculates models tested, successful/failed counts, winner, and average latency.
Spring AI Studio intentionally dispatches model requests in parallel rather than sequentially:
Sequential Execution (Slow):
[ Prompt ] ββ> [ OpenAI (1.8s) ] ββ> [ Anthropic (2.4s) ] ββ> [ Ollama (3.7s) ] ββ> Total: 7.9s
Parallel Execution (Spring AI Studio):
βββ> [ OpenAI (1.8s) ] ββ> Resolves in 1.8s βββ
[ Prompt ] βββΌββ> [ Anthropic (2.4s) ] ββ> Resolves in 2.4s βββΌββ> Benchmark Completed in 3.7s
βββ> [ Ollama (3.7s) ] ββ> Resolves in 3.7s βββ
Application-level response timing depends on:
- Network Round-Trip: Physical distance to cloud datacenters (OpenAI, Anthropic).
- Server Load & Queueing: Cloud provider traffic variations.
- Local Compute Power: Local GPU VRAM bandwidth and quantization speed for Ollama.
- Output Token Length: More verbose explanations require proportionally more generation time.
Determining which model responds first in a concurrent browser environment requires avoiding React stale state closures and race conditions.
// 1. Ref provides synchronous reference across concurrent async callbacks
const firstModelRef = useRef(null);
const [firstModel, setFirstModel] = useState(null);
// 2. In parallel callback handler:
fetchModelResponse(model.id, prompt).then(result => {
const finalStatus = result.error ? 'error' : 'success';
// 3. Atomically lock the first SUCCESSFUL responder
if (finalStatus === 'success' && firstModelRef.current === null) {
firstModelRef.current = model.id;
setFirstModel(model.id);
}
// Update individual card state
setResponses(prev => ({ ...prev, [model.id]: { ... } }));
});- Failed Requests Do Not Win: Immediate errors (e.g.
401 Unauthorizedor connection refused) are ignored. - Single Winner: Once
firstModelRef.currentis set, subsequent responses cannot overwrite it. - Visual Distinction: The winning card receives the gold
.card-firstborder andβ‘ First Responsebadge.
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β BENCHMARK TELEMETRY β
β Models Tested: 3 β Successful: 3 β Failed: 0 β β‘ First Response: OpenAI β Avg: 1.42s β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
The benchmark dashboard derives statistics dynamically from the state of all three model cards:
-
Models Tested: Total registered models (
$N=3$ ). -
Successful Models: Count of cards with
status === 'success'. -
Failed Models: Count of cards with
status === 'error'(hidden if 0). -
First Response: Provider name associated with
firstModelstate. -
Average Response Time: Arithmetic mean of response times across successful responses only:
$$\text{Avg Response Time} = \frac{\sum_{i=1}^{k} \text{Latency}_i}{k} \quad (\text{where } k = \text{Successful Models})$$ Note: Failed models are excluded from the average calculation so missing keys or offline daemons do not distort latency statistics.
Spring-AI-Studio/
βββ Backend/
β βββ .mvn/wrapper/
β β βββ maven-wrapper.jar
β β βββ maven-wrapper.properties # Maven wrapper configuration
β βββ src/
β β βββ main/
β β β βββ java/com/springai/studio/
β β β β βββ AnthropicController.java # REST Controller for Anthropic Claude
β β β β βββ OllamaController.java # REST Controller for Ollama / DeepSeek
β β β β βββ OpenAIController.java # REST Controller for OpenAI GPT-4o
β β β β βββ SpringAiStudioApplication.java# Spring Boot Main Entry Point
β β β βββ resources/
β β β βββ application.properties # Port, API key fallbacks, Ollama model config
β β βββ test/
β β βββ java/com/springai/studio/
β β βββ SpringAiStudioApplicationTests.java # Context loading integration tests
β βββ .env.example # Backend environment variable template
β βββ .gitignore # Git ignore rules for Maven & IDE files
β βββ mvnw # Unix Maven wrapper script
β βββ mvnw.cmd # Windows Maven wrapper script
β βββ pom.xml # Maven dependencies (Spring Boot 3.4.3, Spring AI)
β
βββ Frontend/
β βββ public/
β β βββ branding/
β β β βββ favicon.png # Browser tab icon
β β β βββ spring-ai-studio-logo.png # High-resolution branding logo
β β β βββ spring-ai-studio-wordmark.png # Full horizontal logo wordmark
β βββ src/
β β βββ assets/ # Frontend static assets
β β βββ App.css # Complete UI stylesheet, themes & responsive grid
β β βββ App.jsx # Main workspace, state orchestration & parallel fetch
β β βββ index.css # CSS reset & typography rules
β β βββ main.jsx # React 19 DOM root bootstrap
β βββ .env.example # Frontend environment variable template
β βββ .gitignore # Git ignore rules for node_modules & dist
β βββ eslint.config.js # ESLint configuration
β βββ index.html # HTML5 entry page
β βββ package.json # React 19, Vite 6.2 dependencies & scripts
β βββ package-lock.json # Exact NPM lockfile
β βββ vite.config.js # Vite build configuration with React plugin
β
βββ README.md # Complete project documentation
βββ ...
The backend is built as a modular Spring Boot 3.4.3 application in package com.springai.studio.
Spring Boot Application Context
β
βββ OpenAIController.java
β βββ Injects: OpenAiChatModel ββ> Creates: ChatClient
β βββ Endpoint: POST /api/openai/ask
β
βββ AnthropicController.java
β βββ Injects: AnthropicChatModel ββ> Creates: ChatClient
β βββ Endpoint: POST /api/anthropic/ask
β
βββ OllamaController.java
βββ Injects: OllamaChatModel ββ> Creates: ChatClient
βββ Endpoint: POST /api/ollama/ask
OpenAIController.java: MapsPOST /api/openai/ask, validates payload, callsChatClient, and returns OpenAI GPT-4o output.AnthropicController.java: MapsPOST /api/anthropic/ask, validates payload, callsChatClient, and returns Anthropic Claude output.OllamaController.java: MapsPOST /api/ollama/ask, logs model metadata, extracts text fromChatResponse, and returns DeepSeek output.
The frontend is a modern SPA built with React 19 and Vite 6.2:
- State Management: Uses React standard hooks (
useState,useCallback,useEffect,useRef) for deterministic, lightweight state transitions without heavy Redux/Zustand overhead. - Design Tokens & Theming: CSS custom properties handle light and dark mode styling with smooth color transitions and persistence in
localStorage. - Responsive Layout: CSS Grid and Flexbox dynamically adapt the workspace from single-column mobile views to a 3-column side-by-side comparison matrix on desktop screens.
| Layer | Technology | Verified Version | Purpose / Documentation Link |
|---|---|---|---|
| Language | Java | 21 (LTS) |
Official Java Documentation |
| Backend Framework | Spring Boot | 3.4.3 |
Spring Boot Documentation |
| AI Framework | Spring AI | 1.0.0-M6 |
Spring AI Reference |
| Build Tool | Apache Maven | 3.9+ |
Maven Documentation |
| Frontend Framework | React | 19.0.0 |
React Documentation |
| Frontend Bundler | Vite | 6.2.0 |
Vite Documentation |
| Styling | Vanilla CSS | Modern CSS3 | Custom design system & CSS Grid |
| Cloud AI Model | OpenAI GPT-4o | Cloud API | OpenAI Documentation |
| Cloud AI Model | Anthropic Claude | Cloud API | Anthropic Documentation |
| Local AI Engine | Ollama | deepseek-r1:14b |
Ollama Documentation |
| Frontend Hosting | Vercel | Cloud Edge | Vercel Documentation |
All backend endpoints are stateless HTTP POST handlers accepting JSON payloads and returning plain text strings.
| HTTP Method | Endpoint | Controller | Request Body | Response Body |
|---|---|---|---|---|
POST |
/api/openai/ask |
OpenAIController |
{"prompt": "<string>"} |
Plain text string |
POST |
/api/anthropic/ask |
AnthropicController |
{"prompt": "<string>"} |
Plain text string |
POST |
/api/ollama/ask |
OllamaController |
{"prompt": "<string>"} |
Plain text string |
- URL:
POST /api/openai/ask - Request Headers:
Content-Type: application/json - Request Payload:
{ "prompt": "Explain dependency injection in Spring Boot" } - Success Response (200 OK):
Dependency injection in Spring Boot is a pattern where the Spring IoC container provides required dependencies to classes at runtime, decoupling component creation from business logic. - Error Responses:
400 Bad Request:"Prompt cannot be empty"500 Internal Server Error:"Error from OpenAI: 401 Unauthorized / Invalid API Key"
- URL:
POST /api/anthropic/ask - Request Payload:
{ "prompt": "Compare REST vs GraphQL" } - Success Response (200 OK):
REST uses fixed endpoints and standard HTTP methods, whereas GraphQL allows clients to request exact fields from a single endpoint.
- URL:
POST /api/ollama/ask - Request Payload:
{ "prompt": "Write a Java Stream example" } - Success Response (200 OK):
List<Integer> evens = numbers.stream().filter(n -> n % 2 == 0).toList(); - Error Responses:
500 Internal Server Error:"Error from Ollama: Connection refused"
| Variable Name | Default / Fallback | Description | Required? |
|---|---|---|---|
SPRING_AI_OPENAI_API_KEY |
dummy-openai-key |
OpenAI API Secret Key (sk-...). |
Optional (for OpenAI) |
SPRING_AI_ANTHROPIC_API_KEY |
dummy-anthropic-key |
Anthropic API Secret Key (sk-ant-...). |
Optional (for Anthropic) |
SPRING_AI_OLLAMA_BASE_URL |
http://localhost:11434 |
Ollama daemon base URL. | Optional (for Local LLM) |
SPRING_AI_OLLAMA_MODEL |
deepseek-r1:14b |
Tag name of the installed Ollama model. | Optional (for Local LLM) |
server.port |
8080 |
Spring Boot HTTP listening port. | Built-in |
| Variable Name | Local Development Value | Production Example | Description |
|---|---|---|---|
VITE_API_BASE_URL |
http://localhost:8080 |
https://api.yourdomain.com |
Base URL pointing to the Spring Boot REST API. |
Caution
Security Rules:
- Never commit
.envfiles containing live API keys. - Never put OpenAI/Anthropic secret keys into the frontend
.env. - In production, update
VITE_API_BASE_URLto point to your deployed backend URL.
Follow these exact steps to run the application locally on Windows or Unix:
git clone https://github.com/Mohammad-Asfin/Spring-AI-Studio.git
cd Spring-AI-Studio# Windows (PowerShell)
$env:SPRING_AI_OPENAI_API_KEY="sk-your-openai-api-key"
$env:SPRING_AI_ANTHROPIC_API_KEY="sk-ant-your-anthropic-api-key"
# Linux / macOS (Bash)
export SPRING_AI_OPENAI_API_KEY="sk-your-openai-api-key"
export SPRING_AI_ANTHROPIC_API_KEY="sk-ant-your-anthropic-api-key"# Navigate to Backend
cd "d:\Java Full Stack\Spring AI\Backend"
# Build and start Spring Boot
mvn spring-boot:runBackend initializes on http://localhost:8080.
# Navigate to Frontend
cd "d:\Java Full Stack\Spring AI\Frontend"
# Install dependencies
npm install
# Start Vite dev server
npm run devFrontend initializes on http://localhost:5173. Open http://localhost:5173 in your browser.
- Obtain an API key from platform.openai.com/api-keys.
- Set the environment variable
SPRING_AI_OPENAI_API_KEY. - If omitted, the backend falls back to
dummy-openai-keyto allow Spring context loading; the UI card will display diagnostic guidance.
- Obtain an API key from console.anthropic.com.
- Set the environment variable
SPRING_AI_ANTHROPIC_API_KEY. - Missing keys display an actionable card error.
- Download and install Ollama from ollama.com.
- Pull the configured model in your terminal:
ollama pull deepseek-r1:14b
- Start the Ollama daemon:
ollama serve
- Verify accessibility at
http://localhost:11434.
Important
Current Repository Status:
- Frontend: Deployed live on Vercel (
https://spring-ai-studio-psi.vercel.app/). - Backend: Backend deployment is not currently configured in this repository.
To run the backend in a production cloud environment, deploy Spring Boot to a Java 21 runtime platform (such as Railway, Render, AWS, or Docker).
- Create an account on Railway.app.
- Click New Project β Deploy from GitHub repo.
- Select
Mohammad-Asfin/Spring-AI-Studio. - In Project Settings:
- Root Directory: Set to
Backend. - Build Command:
mvn clean package -DskipTests. - Start Command:
java -jar target/*.jar.
- Root Directory: Set to
- Under Variables, add
SPRING_AI_OPENAI_API_KEYandSPRING_AI_ANTHROPIC_API_KEY. - Click Generate Public Domain to obtain your backend URL (e.g.
https://spring-ai-backend.up.railway.app).
(Recommended Future Improvement: Add a Dockerfile to the Backend/ directory)
# Stage 1: Build JAR using Maven and JDK 21
FROM eclipse-temurin:21-jdk-alpine AS build
WORKDIR /app
COPY pom.xml .
COPY src ./src
COPY .mvn ./.mvn
COPY mvnw .
RUN ./mvnw clean package -DskipTests
# Stage 2: Minimal JRE Runtime
FROM eclipse-temurin:21-jre-alpine
WORKDIR /app
COPY --from=build /app/target/*.jar app.jar
EXPOSE 8080
ENV PORT=8080
ENTRYPOINT ["java", "-jar", "app.jar"]The React single-page application is deployed live at:
https://spring-ai-studio-psi.vercel.app/
- Log in to Vercel and click Add New Project.
- Select your
Spring-AI-StudioGitHub repository. - Configure Project Settings:
- Framework Preset:
Vite - Root Directory:
Frontend - Build Command:
npm run build - Output Directory:
dist
- Framework Preset:
- Under Environment Variables, add:
VITE_API_BASE_URL: The public URL of your deployed backend (e.g.https://your-backend-domain.com).
- Click Deploy.
To connect the deployed frontend with the deployed backend:
- Deploy Backend: Deploy Spring Boot on Railway or Render and obtain the public HTTPS URL.
- Update Frontend Environment Variable: In Vercel Project Settings β Environment Variables, set:
VITE_API_BASE_URL=https://your-deployed-backend-url.com
- Trigger Vercel Redeploy: Redeploy the frontend so Vite compiles the updated API base URL into the bundle.
- Configure CORS: Ensure the backend allows requests from your Vercel domain.
| Security Aspect | Current Codebase Implementation | Recommended Production Hardening |
|---|---|---|
| API Key Isolation | β Keys stored in backend environment variables only. Zero keys in client bundles. | Use cloud secret managers (AWS Secrets Manager, Railway Secrets). |
| CORS Policy | @CrossOrigin("*") on all controllers for seamless local multi-port development. |
Restrict origins: @CrossOrigin(origins = "https://spring-ai-studio-psi.vercel.app"). |
| Error Masking | Returns e.getMessage() for debugging. |
Sanitize internal stack traces before returning HTTP 500 to clients. |
| Transport Security | HTTP on localhost. |
Enforce HTTPS via TLS termination at the edge/load balancer. |
cd Frontend
npm run buildExpected: β built in ~1.0s producing dist/ bundle.
cd Backend
mvn testExpected: BUILD SUCCESS with context loading verified.
| Level | Scope | Status in CI / Clean Environment |
|---|---|---|
| Build Tested | Maven compile, test context load, Vite bundle compilation. | β Verified (100% automated passing) |
| Live Provider Tested | Live OpenAI, Anthropic, or Ollama round-trips. | π Requires valid live API keys & local daemon |
| Symptom | Probable Cause | Diagnostic & Solution |
|---|---|---|
Backend fails on startup (Port 8080 already in use) |
Another process is occupying port 8080. | Change server.port=8081 in application.properties and update VITE_API_BASE_URL=http://localhost:8081. |
OpenAI card displays Failed / 401 Unauthorized |
Missing or expired OpenAI API key. | Set SPRING_AI_OPENAI_API_KEY="sk-..." in environment before starting the backend. |
Anthropic card displays Failed |
Missing or invalid Anthropic API key. | Set SPRING_AI_ANTHROPIC_API_KEY="sk-ant-..." in environment. |
Ollama card displays Requires Ollama |
Ollama daemon is not running on port 11434. | Open terminal, run ollama serve, and verify http://localhost:11434. |
Ollama returns model not found |
The deepseek-r1:14b model has not been downloaded. |
Run ollama pull deepseek-r1:14b in terminal. |
Frontend displays Failed to fetch / Network Error |
Backend is offline or blocked by CORS. | Ensure backend is active on 8080 and check browser DevTools Network tab. |
Windows Maven Wrapper error ('C:\Users\MD' is not recognized) |
Space in Windows user folder path. | Run mvn spring-boot:run directly instead of ./mvnw. |
Try these prompts to benchmark reasoning styles across models:
- System Architecture:
"Explain the difference between Dependency Injection and Inversion of Control in Spring Boot with a concrete code snippet."
- API Design Trade-offs:
"Compare REST APIs with GraphQL across performance, over-fetching, caching, and client flexibility."
- Modern Java Syntax:
"Write a Java 21 Stream pipeline to group a list of transactions by currency and calculate the total sum for each."
- Spring AI Internals:
"How does the Spring AI ChatClient abstraction simplify switching between OpenAI and Anthropic compared to raw HTTP clients?"
Application-level response timing depends on:
- Model Architecture: Frontier models (
GPT-4o) balance reasoning depth and speed; dense reasoning models (deepseek-r1) perform step-by-step chain-of-thought token generation. - Local Hardware Constraints: Local Ollama token generation speed depends directly on available GPU VRAM bandwidth.
- Geographic Network Latency: Cloud API response times include TLS handshakes and physical network hops.
Tip: To add additional screenshots of the model comparison matrix or dark mode, save images to
Frontend/public/branding/and embed them here.
- Streaming Token Generation: Implement Server-Sent Events (SSE) / WebSockets using Spring AI reactive streaming.
- Dynamic Model Selector: UI dropdown to select alternate models (e.g.
gpt-4o-mini,claude-3-5-haiku,llama3.3). - Token Count & Cost Tracking: Live estimation of input/output token usage and approximate API cost per prompt.
- Health & Actuator Endpoint: Add
spring-boot-starter-actuatorwith/actuator/healthfor cloud deployment monitoring. - Benchmark Export: Export comparison sessions and timing metrics to JSON, CSV, or Markdown summaries.
- Containerization: Commit official multi-stage
Dockerfileanddocker-compose.yml.
Contributions are welcome! Please follow these steps:
- Fork the Repository on GitHub.
- Create a Feature Branch:
git checkout -b feature/streaming-support
- Commit Your Changes:
git commit -m "feat: add token streaming via SSE" - Push to Your Branch:
git push origin feature/streaming-support
- Open a Pull Request describing your additions.
This project is open-source. Feel free to use, modify, and distribute for educational, research, and commercial purposes with attribution.
- Live Frontend Application: https://spring-ai-studio-psi.vercel.app/
- GitHub Repository: https://github.com/Mohammad-Asfin/Spring-AI-Studio
- Spring AI Official Documentation: https://docs.spring.io/spring-ai/reference/
- Spring Boot 3.4 Documentation: https://docs.spring.io/spring-boot/index.html
- React 19 Documentation: https://react.dev/
- Vite Documentation: https://vite.dev/
- Ollama Documentation: https://ollama.com/
- OpenAI Platform Documentation: https://platform.openai.com/docs/
- Anthropic Documentation: https://docs.anthropic.com/
- Railway Spring Boot Deployment Guide: https://docs.railway.com/guides/spring-boot
- Render Docker Deployment Guide: https://render.com/docs/docker
Spring AI Studio β’ Developed by Mohammad Asfin
