Project Name: File Processing & Conversion System
Goal: Build an industry-ready document and file processing platform similar in capability to online PDF/file tools such as iLovePDF.
The system will support:
- PDF manipulation
- PDF conversion
- Office document conversion
- Image processing
- OCR
- File upload/download
- Background processing
- Temporary file lifecycle
- Cloud object storage
- Authentication and user history
- Scalable worker processing
Layer Technology
Frontend Angular Backend ASP.NET Core Web API / .NET 10 Architecture Clean Architecture Database SQL Server ORM Entity Framework Core Cache Redis Background Jobs Hangfire Object Storage Azure Blob Storage / Amazon S3 Local Dev Storage MinIO PDF Generation QuestPDF PDF Reading/Extraction PdfPig PDF Compression/Optimization Ghostscript Office → PDF LibreOffice Image Processing ImageSharp Advanced Image Processing Magick.NET OCR Tesseract Logging Serilog Validation FluentValidation Testing xUnit + integration tests API Documentation OpenAPI / Swagger Containerization Docker + Docker Compose Reverse Proxy Nginx
The application should not be designed as a collection of controllers that directly call third-party libraries.
The architecture should be:
Angular
|
v
ASP.NET Core API
|
v
Application
|
+------------------+
| |
v v
Domain Infrastructure
|
+-------------+-------------+
| | |
v v v
PDF Engine Office Engine Image/OCR
| | |
+-------------+-------------+
|
v
Object Storage
|
v
SQL Server
The main principle is:
Application defines what the system needs. Infrastructure defines how external libraries/tools perform it.
For example:
Application
|
+-- IPdfCompressionService
^
|
Infrastructure
|
+-- GhostscriptPdfCompressionService
The Application layer must not depend directly on Ghostscript.
The system follows a Modular Monolith architecture.
The application is deployed as one backend application initially, but the business capabilities are isolated into independent modules. Each module owns its own Domain, Application and module-specific Infrastructure.
Cross-cutting technical capabilities such as database access, object storage, Redis, Hangfire and logging are placed in Shared Infrastructure.
EF Core migrations are maintained in a separate Migration project.
FileProcessingSystem/
│
├── application/
│ │
│ ├── API/
│ │ └── FileProcessingSystem.API/
│ │
│ ├── Modules/
│ │ │
│ │ ├── FileManagement/
│ │ │ ├── Domain/
│ │ │ │ ├── Entities/
│ │ │ │ ├── Enums/
│ │ │ │ └── ValueObjects/
│ │ │ │
│ │ │ ├── Application/
│ │ │ │ ├── Abstractions/
│ │ │ │ ├── Features/
│ │ │ │ ├── DTOs/
│ │ │ │ └── Validators/
│ │ │ │
│ │ │ └── Infrastructure/
│ │ │ ├── Persistence/
│ │ │ └── Services/
│ │ │
│ │ ├── PDF/
│ │ │ ├── Domain/
│ │ │ ├── Application/
│ │ │ │ ├── Abstractions/
│ │ │ │ ├── Features/
│ │ │ │ ├── DTOs/
│ │ │ │ └── Validators/
│ │ │ └── Infrastructure/
│ │ │ ├── PdfPig/
│ │ │ ├── QuestPDF/
│ │ │ └── Ghostscript/
│ │ │
│ │ ├── Conversion/
│ │ │ ├── Domain/
│ │ │ ├── Application/
│ │ │ │ ├── Abstractions/
│ │ │ │ ├── Features/
│ │ │ │ ├── DTOs/
│ │ │ │ └── Validators/
│ │ │ └── Infrastructure/
│ │ │ └── LibreOffice/
│ │ │
│ │ ├── ImageProcessing/
│ │ │ ├── Domain/
│ │ │ ├── Application/
│ │ │ │ ├── Abstractions/
│ │ │ │ ├── Features/
│ │ │ │ ├── DTOs/
│ │ │ │ └── Validators/
│ │ │ └── Infrastructure/
│ │ │ ├── ImageSharp/
│ │ │ └── MagickNET/
│ │ │
│ │ └── OCR/
│ │ ├── Domain/
│ │ ├── Application/
│ │ │ ├── Abstractions/
│ │ │ ├── Features/
│ │ │ ├── DTOs/
│ │ │ └── Validators/
│ │ └── Infrastructure/
│ │ └── Tesseract/
│ │
│ ├── Shared/
│ │ │
│ │ ├── CommonLibrary/
│ │ │ ├── Results/
│ │ │ ├── Exceptions/
│ │ │ ├── Extensions/
│ │ │ ├── Constants/
│ │ │ └── Helpers/
│ │ │
│ │ └── Infrastructure/
│ │ ├── Persistence/
│ │ ├── Storage/
│ │ │ ├── Local/
│ │ │ ├── MinIO/
│ │ │ ├── AzureBlob/
│ │ │ └── S3/
│ │ ├── Redis/
│ │ ├── Hangfire/
│ │ ├── Logging/
│ │ └── Security/
│ │
│ └── Migration/
│ └── FileProcessingSystem.Migration/
│ ├── Migrations/
│ └── Program.cs
│
├── frontend/
│ └── file-processing-system-ui/
│
├── tests/
│ ├── FileProcessingSystem.UnitTests/
│ └── FileProcessingSystem.IntegrationTests/
│
├── docker/
│ ├── api/
│ ├── worker/
│ └── nginx/
│
├── docs/
│ ├── architecture/
│ ├── api/
│ └── deployment/
│
├── docker-compose.yml
├── Directory.Build.props
├── Directory.Packages.props
└── README.md
A Modular Monolith gives the project strong module boundaries without the operational complexity of microservices.
File Processing System
│
┌───────────┴───────────┐
│ ASP.NET Core │
│ Application │
└───────────┬───────────┘
│
┌────────────────────┼────────────────────┐
│ │ │
▼ ▼ ▼
File Management PDF Conversion
│ │ │
└──────────────┬─────┴──────────────┬─────┘
│ │
▼ ▼
Image Processing OCR
│
└──────────┬──────────┘
│
▼
Shared Infrastructure
Each module has a clear responsibility and should communicate through application contracts rather than directly accessing another module's internal implementation.
- Easier development than microservices
- One deployment unit
- Simple local development
- Clear business boundaries
- Independent module ownership
- Easier testing
- Easier future extraction into microservices if required
- Shared infrastructure without duplicating technical code
Responsibility:
Upload
Download
Delete
File validation
File metadata
File lifecycle
File ownership
Temporary file management
Structure:
FileManagement/
├── Domain/
├── Application/
└── Infrastructure/
This module owns file-related business rules.
It should not know how Azure Blob, S3 or MinIO works internally.
It uses:
IFileStorageServicewhich is provided by Shared Infrastructure.
Responsibility:
Merge
Split
Remove pages
Extract pages
Reorder pages
Rotate
Watermark
Page numbers
Protect
Unlock with valid credentials
Compress
Extract text
Render pages
PDF metadata
Structure:
PDF/
├── Domain/
├── Application/
│ ├── Abstractions/
│ ├── Features/
│ ├── DTOs/
│ └── Validators/
│
└── Infrastructure/
├── PdfPig/
├── QuestPDF/
└── Ghostscript/
The PDF module owns the PDF business/use-case logic.
Third-party PDF libraries stay inside:
PDF.Infrastructure
Responsibility:
Word → PDF
Excel → PDF
PowerPoint → PDF
PDF → Word
PDF → Excel
PDF → PowerPoint
HTML → PDF
Structure:
Conversion/
├── Domain/
├── Application/
└── Infrastructure/
└── LibreOffice/
The Application layer defines:
public interface IOfficeToPdfService
{
Task<Stream> ConvertAsync(
Stream input,
string fileName,
CancellationToken cancellationToken);
}Infrastructure implements it using LibreOffice.
Responsibility:
Resize
Compress
Crop
Rotate
Convert
Watermark
Thumbnail
Image → PDF
PDF → Image where applicable
Structure:
ImageProcessing/
├── Domain/
├── Application/
└── Infrastructure/
├── ImageSharp/
└── MagickNET/
Use ImageSharp as the normal .NET image-processing engine.
Use Magick.NET only where advanced format support or processing requirements justify it.
Responsibility:
Image → Text
PDF → Text
Scanned PDF → Searchable PDF
OCR processing
Language configuration
OCR result handling
Structure:
OCR/
├── Domain/
├── Application/
└── Infrastructure/
└── Tesseract/
The OCR engine itself remains an infrastructure concern.
Shared Infrastructure contains technical services that are used by multiple modules.
Shared/
└── Infrastructure/
├── Persistence/
├── Storage/
├── Redis/
├── Hangfire/
├── Logging/
└── Security/
Contains:
EF Core
SQL Server
DbContext
Common persistence configuration
Transaction infrastructure
The shared persistence layer should provide the technical database infrastructure.
Module-specific entities/configurations should remain associated with their respective module.
Recommended direction:
Module Domain
↓
Shared Persistence
↓
SQL Server
Shared.Infrastructure/
└── Storage/
├── Local/
├── MinIO/
├── AzureBlob/
└── S3/
Common abstraction:
public interface IFileStorageService
{
Task<StoredFileResult> UploadAsync(
Stream stream,
string fileName,
string contentType,
CancellationToken cancellationToken);
Task<Stream> OpenReadAsync(
string storageKey,
CancellationToken cancellationToken);
Task DeleteAsync(
string storageKey,
CancellationToken cancellationToken);
}Modules use the abstraction.
They do not directly reference:
Azure.Storage.Blobs
Amazon.S3
MinIO SDK
Shared Redis infrastructure handles:
Caching
Rate limiting state
Short-lived job state
Distributed coordination
It must not be used as the primary storage for uploaded files.
SQL Server
↓
Persistent metadata
Redis
↓
Temporary/cache state
Object Storage
↓
Actual files
Hangfire is shared because multiple modules can create processing jobs.
PDF Module
│
Conversion Module
│
Image Module
│
OCR Module
│
▼
Shared Hangfire Infrastructure
│
▼
Background Worker
Example:
BackgroundJob.Enqueue(() =>
pdfCompressionService.CompressAsync(jobId));The actual processing implementation remains inside the owning module.
Use:
Serilog
as shared infrastructure.
Every module should produce structured logs containing useful context such as:
CorrelationId
JobId
UserId
Module
Operation
Duration
Status
Error
EF Core migrations should be isolated from the runtime API and Infrastructure projects.
application/
└── Migration/
└── FileProcessingSystem.Migration/
├── Migrations/
└── Program.cs
Purpose:
Runtime application
≠
Database migration tooling
The migration project references the required Infrastructure and DbContext components but is not part of the API request-processing pipeline.
Typical workflow:
Change Entity
↓
Create Migration
↓
Migration Project
↓
SQL Server
Example:
dotnet ef migrations add InitialCreate \
--project application/Migration/FileProcessingSystem.Migration \
--startup-project application/API/FileProcessingSystem.APIDatabase update:
dotnet ef database update \
--project application/Migration/FileProcessingSystem.Migration \
--startup-project application/API/FileProcessingSystem.APIThe modular monolith must enforce dependency direction.
API
│
▼
Application
│
▼
Domain
Infrastructure implements Application abstractions:
Application
↑
│
Infrastructure
Modules should not depend on another module's Infrastructure.
PDF.Application
↓
PDF abstraction
PDF.Infrastructure
↓
PDF third-party library
PDF.Application
↓
Conversion.Infrastructure
↓
LibreOffice
Instead, cross-module communication should use contracts/application abstractions.
When one module needs another module:
PDF
│
▼
Application Contract
│
▼
Other Module
Do not access:
OtherModule.Infrastructure
OtherModule.DbContext
OtherModule.Repository
OtherModule.InternalService
Example:
Conversion Module
|
| needs file
v
IFileStorageService
|
v
Shared Infrastructure
|
v
Azure Blob / S3 / MinIO
For a PDF compression request:
Angular
|
| POST /api/pdf/compress
v
API Controller
|
v
PDF Application
|
v
PDF Compression Use Case
|
+------> File Management / Storage abstraction
|
+------> Shared Hangfire
|
v
PDF Infrastructure
|
v
Ghostscript
|
v
Output Stream
|
v
Shared Storage
|
v
Azure Blob / S3 / MinIO
|
v
SQL Server metadata
The API never directly invokes Ghostscript.
Angular remains a separate frontend application:
frontend/
└── file-processing-system-ui/
Recommended:
src/app/
├── core/
│ ├── auth/
│ ├── guards/
│ ├── interceptors/
│ ├── services/
│ └── models/
│
├── shared/
│ ├── components/
│ ├── directives/
│ ├── pipes/
│ └── models/
│
├── features/
│ ├── home/
│ ├── file-management/
│ ├── pdf/
│ │ ├── merge/
│ │ ├── split/
│ │ ├── compress/
│ │ ├── rotate/
│ │ └── watermark/
│ │
│ ├── conversion/
│ ├── images/
│ ├── ocr/
│ ├── jobs/
│ └── account/
│
├── app.routes.ts
└── app.config.ts
Angular feature boundaries should mirror backend business modules where practical.
Angular PDF Feature
↓
PDF API
↓
PDF Module
Use a single SQL Server database initially.
SQL Server
|
+-- File Management tables
+-- PDF/processing tables
+-- Conversion/job tables
+-- User/account tables
+-- Audit tables
Recommended core tables:
Users
Files
FileVersions
StorageObjects
ProcessingJobs
ProcessingJobFiles
UserOperationHistory
AuditLogs
The database stores metadata.
Large binary files are stored in object storage.
SQL Server
-------------------------
FileId
UserId
StorageKey
FileName
Size
ContentType
Status
CreatedAt
ExpiresAt
+
Object Storage
-------------------------
actual PDF/DOCX/XLSX/JPG/etc.
MinIO
Choose one managed provider:
Azure Blob Storage
or:
Amazon S3
Recommended Azure-oriented deployment:
Angular
↓
Azure Static Web Apps / CDN
↓
ASP.NET Core API
↓
Background Workers
↓
Azure Blob Storage
↓
Azure SQL
↓
Redis
The storage abstraction allows the provider to change without changing module business logic.
┌─────────────────────────────────────────────────────────────┐
│ Angular Frontend │
│ │
│ File | PDF | Conversion | Image | OCR | Jobs | Account │
└─────────────────────────────┬───────────────────────────────┘
│ HTTPS
▼
┌─────────────────────────────────────────────────────────────┐
│ ASP.NET Core API │
│ │
│ Controllers | Auth | Middleware | OpenAPI | DI │
└─────────────────────────────┬───────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ MODULAR MONOLITH │
│ │
│ ┌────────────┐ ┌────────┐ ┌────────────┐ │
│ │ File │ │ PDF │ │ Conversion │ │
│ │ Management │ │ Module │ │ Module │ │
│ └────────────┘ └────────┘ └────────────┘ │
│ │
│ ┌────────────────┐ ┌────────────┐ │
│ │ Image │ │ OCR │ │
│ │ Processing │ │ Module │ │
│ └────────────────┘ └────────────┘ │
└─────────────────────────────┬───────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ SHARED INFRASTRUCTURE │
│ │
│ EF Core | Storage | Redis | Hangfire | Serilog | Security │
└──────────────┬────────────────┬────────────────┬────────────┘
│ │ │
▼ ▼ ▼
SQL Server Redis/Queue Blob/S3/MinIO
│ │
│ │
│ Actual Files
│
Metadata/State
MODULAR MONOLITH
│
┌───────────────┼────────────────┐
│ │ │
▼ ▼ ▼
Modules Shared Infrastructure Migration
│ │ │
│ │ └── Separate project
│ │
│ ├── SQL Server
│ ├── Storage
│ ├── Redis
│ ├── Hangfire
│ └── Serilog
│
├── File Management
├── PDF
├── Conversion
├── Image Processing
└── OCR
This structure keeps the application as a single deployable monolith today, while maintaining clear module boundaries so individual modules can be extracted into separate services later if scale or organizational requirements justify microservices.
Project:
FileProcessingSystem.API
Responsibilities:
- HTTP endpoints
- Controllers
- Authentication configuration
- Middleware
- Exception handling
- API versioning
- Swagger/OpenAPI
- Dependency Injection registration
- Request/response models where appropriate
Example:
API/
├── Controllers/
│ ├── FilesController.cs
│ ├── PdfController.cs
│ ├── ConversionController.cs
│ ├── ImageController.cs
│ ├── OcrController.cs
│ └── JobsController.cs
│
├── Middleware/
│ ├── ExceptionHandlingMiddleware.cs
│ └── CorrelationIdMiddleware.cs
│
├── Extensions/
│ ├── ServiceCollectionExtensions.cs
│ └── ApplicationBuilderExtensions.cs
│
└── Program.cs
Controller example:
[ApiController]
[Route("api/pdf")]
public sealed class PdfController : ControllerBase
{
private readonly IPdfMergeService _pdfMergeService;
public PdfController(IPdfMergeService pdfMergeService)
{
_pdfMergeService = pdfMergeService;
}
[HttpPost("merge")]
public async Task<IActionResult> Merge(
[FromForm] List<IFormFile> files,
CancellationToken cancellationToken)
{
var result = await _pdfMergeService
.MergeAsync(files, cancellationToken);
return File(
result.Content,
result.ContentType,
result.FileName);
}
}The controller should remain thin.
Project:
FileProcessingSystem.Domain
Domain contains business concepts that should not depend on ASP.NET, EF Core, Redis, cloud SDKs or PDF libraries.
Recommended:
Domain/
├── Entities/
│ ├── FileRecord.cs
│ ├── ProcessingJob.cs
│ ├── ProcessingOperation.cs
│ ├── UserFile.cs
│ └── StoredFile.cs
│
├── Enums/
│ ├── FileStatus.cs
│ ├── JobStatus.cs
│ ├── FileOperationType.cs
│ └── StorageProvider.cs
│
├── ValueObjects/
│ ├── FileSize.cs
│ └── FileIdentifier.cs
│
└── Constants/
Example:
public enum ProcessingJobStatus
{
Queued,
Processing,
Completed,
Failed,
Cancelled
}Project:
FileProcessingSystem.Application
This is the main business/use-case layer.
Recommended structure:
Application/
├── Abstractions/
│ ├── Storage/
│ │ └── IFileStorageService.cs
│ │
│ ├── PDF/
│ │ ├── IPdfMergeService.cs
│ │ ├── IPdfSplitService.cs
│ │ ├── IPdfCompressionService.cs
│ │ ├── IPdfPageService.cs
│ │ ├── IPdfWatermarkService.cs
│ │ ├── IPdfProtectionService.cs
│ │ ├── IPdfRenderingService.cs
│ │ └── IPdfTextExtractionService.cs
│ │
│ ├── Conversion/
│ │ ├── IOfficeToPdfService.cs
│ │ ├── IPdfToWordService.cs
│ │ ├── IPdfToExcelService.cs
│ │ └── IPdfToPowerPointService.cs
│ │
│ ├── Image/
│ │ └── IImageProcessingService.cs
│ │
│ ├── OCR/
│ │ └── IOcrService.cs
│ │
│ └── Jobs/
│ └── IProcessingJobService.cs
│
├── Features/
│ ├── Files/
│ ├── PDF/
│ ├── Conversion/
│ ├── Images/
│ ├── OCR/
│ └── Jobs/
│
├── DTOs/
├── Validators/
└── DependencyInjection.cs
Project:
FileProcessingSystem.CommonLibrary
This project should contain genuinely shared technical utilities.
Recommended:
CommonLibrary/
├── Results/
├── Exceptions/
├── Extensions/
├── Constants/
├── Helpers/
├── Security/
├── Serialization/
└── Time/
Examples:
Result<T>
Error
PagedResult<T>
DateTime extensions
String extensions
File-name sanitization
Common exception types
Do not turn CommonLibrary into a dumping ground.
Project:
FileProcessingSystem.Infrastructure
This project contains all external implementations.
Infrastructure/
├── Persistence/
│ ├── AppDbContext.cs
│ ├── Configurations/
│ └── Repositories/
│
├── Storage/
│ ├── Local/
│ ├── AzureBlob/
│ ├── S3/
│ └── Minio/
│
├── PDF/
│ ├── PdfPig/
│ ├── QuestPdf/
│ └── Ghostscript/
│
├── Office/
│ └── LibreOffice/
│
├── Images/
│ ├── ImageSharp/
│ └── MagickNet/
│
├── OCR/
│ └── Tesseract/
│
├── BackgroundJobs/
│ └── Hangfire/
│
├── Caching/
│ └── Redis/
│
├── Logging/
│ └── Serilog/
│
└── DependencyInjection.cs
Project:
FileProcessingSystem.Infrastructure.Migration
Keep EF Core migrations separate from the runtime Infrastructure project.
Migration/
├── Migrations/
│ ├── 202608290001_InitialCreate.cs
│ ├── 202608290001_InitialCreate.Designer.cs
│ └── AppDbContextModelSnapshot.cs
│
├── Program.cs
└── FileProcessingSystem.Infrastructure.Migration.csproj
This follows the existing project approach where the migration project is under:
Infrastructure/
└── Migration/
Project:
frontend/file-processing-system-ui
Recommended Angular structure:
src/
├── app/
│ │
│ ├── core/
│ │ ├── auth/
│ │ ├── guards/
│ │ ├── interceptors/
│ │ ├── services/
│ │ └── models/
│ │
│ ├── shared/
│ │ ├── components/
│ │ ├── directives/
│ │ ├── pipes/
│ │ └── models/
│ │
│ ├── features/
│ │ ├── home/
│ │ ├── pdf/
│ │ │ ├── merge/
│ │ │ ├── split/
│ │ │ ├── compress/
│ │ │ ├── rotate/
│ │ │ ├── watermark/
│ │ │ └── protect/
│ │ │
│ │ ├── conversion/
│ │ ├── images/
│ │ ├── ocr/
│ │ ├── jobs/
│ │ └── account/
│ │
│ ├── app.routes.ts
│ └── app.config.ts
│
└── assets/
Angular should communicate only with the ASP.NET API.
Angular Service
|
v
HTTP API
|
v
ASP.NET Controller
The standard pipeline should be:
User
|
v
Angular
|
| multipart/form-data
v
ASP.NET API
|
v
File Validation
|
+---- Invalid ---> 400
|
v
Object Storage
|
v
Create Processing Job
|
v
Hangfire
|
v
Worker
|
v
Processing Engine
|
v
Output Object Storage
|
v
Update Job
|
v
Angular
|
v
Download
Do not couple the application to a physical disk.
Create:
public interface IFileStorageService
{
Task<StoredFileResult> UploadAsync(
Stream stream,
string fileName,
string contentType,
CancellationToken cancellationToken);
Task<Stream> OpenReadAsync(
string storageKey,
CancellationToken cancellationToken);
Task DeleteAsync(
string storageKey,
CancellationToken cancellationToken);
}Implementations:
IFileStorageService
|
+-- LocalFileStorageService
+-- MinioFileStorageService
+-- AzureBlobStorageService
+-- S3FileStorageService
Use:
MinIO
because it provides S3-compatible object storage locally.
Recommended primary choice:
Azure Blob Storage
or:
Amazon S3
depending on deployment ecosystem.
The Application layer should not know which provider is being used.
SQL Server stores metadata and application state.
Do not use SQL Server as the default place for large uploaded files.
Recommended database:
Users
Files
FileVersions
ProcessingJobs
ProcessingJobFiles
Operations
UserOperationHistory
StorageObjects
ApiKeys
RefreshTokens
AuditLogs
Example:
Files
--------------------------------
Id
UserId
OriginalFileName
StoredFileName
ContentType
Extension
SizeInBytes
StorageKey
StorageProvider
Status
CreatedAt
ExpiresAt
The actual file:
Azure Blob / S3 / MinIO
The database stores:
storageKey
Use for:
- Text extraction
- Reading PDF structure
- Page inspection
- PDF metadata
- Text analysis
Example:
using UglyToad.PdfPig;
using var document = PdfDocument.Open(stream);
foreach (var page in document.GetPages())
{
var text = page.Text;
}Do not use PdfPig as the only PDF engine for every operation.
Use for:
- Creating new PDFs
- Reports
- Generated documents
- Images to PDF
- Application-generated PDF output
Example:
Document.Create(document =>
{
document.Page(page =>
{
page.Content()
.Text("Generated PDF");
});
})
.GeneratePdf(outputStream);Responsibility:
QuestPDF
=
PDF Generation
Not:
QuestPDF
=
Every PDF manipulation operation
Use for:
- PDF compression
- PDF optimization
- PDF rendering-related workflows where appropriate
- PDF compatibility transformations
Architecture:
IPdfCompressionService
|
v
GhostscriptPdfCompressionService
|
v
Ghostscript executable
|
v
Optimized PDF
The Ghostscript process should execute inside a controlled worker environment.
Never allow arbitrary command-line arguments from the user to reach the process.
Use for:
DOC
DOCX
XLS
XLSX
PPT
PPTX
|
v
LibreOffice headless
|
v
PDF
Architecture:
IOfficeToPdfService
|
v
LibreOfficeService
Run LibreOffice only in the worker/container environment.
The API should never directly expose operating-system process execution to the client.
Use ImageSharp as the default .NET image processing library.
Operations:
Resize
Crop
Rotate
Compress
Format conversion
Thumbnail
Watermark
Example:
using var image = await Image.LoadAsync(inputStream);
image.Mutate(x =>
x.Resize(1200, 0));
await image.SaveAsync(outputStream, cancellationToken);Use when ImageSharp does not provide the required advanced format or operation.
Typical responsibilities:
Advanced image formats
Advanced transformations
Image metadata handling
Specialized conversion
Do not automatically send every image through both ImageSharp and Magick.NET.
Recommended engine:
Tesseract
Pipeline:
Scanned PDF
|
v
PDF Renderer
|
v
Page Image
|
v
Tesseract
|
v
Recognized Text
For searchable PDF:
Scanned PDF
|
v
Render pages
|
v
OCR
|
v
Text + original page image
|
v
Searchable PDF
Large operations should not execute inside a normal HTTP request.
Use:
ASP.NET Core
|
v
Create Job
|
v
Hangfire
|
v
Worker
|
v
Processing
Example:
BackgroundJob.Enqueue(() =>
pdfCompressionService.CompressAsync(jobId));Job states:
Queued
↓
Processing
↓
Completed
Failure:
Processing
↓
Failed
The API should return a job identifier for long-running work.
Redis should be used for:
- Distributed cache
- Rate limiting state
- Temporary job progress
- Short-lived data
- Distributed coordination where required
Do not store large uploaded PDFs in Redis.
Correct separation:
SQL Server
↓
Persistent metadata
Redis
↓
Temporary/cache data
Blob/S3
↓
Actual files
Recommended endpoints:
/api/files
/api/jobs
/api/pdf/merge
/api/pdf/split
/api/pdf/compress
/api/pdf/rotate
/api/pdf/watermark
/api/pdf/protect
/api/pdf/unlock
/api/pdf/extract-pages
/api/pdf/remove-pages
/api/pdf/extract-text
/api/pdf/render
/api/conversion/word-to-pdf
/api/conversion/excel-to-pdf
/api/conversion/powerpoint-to-pdf
/api/conversion/pdf-to-word
/api/conversion/pdf-to-excel
/api/images/resize
/api/images/compress
/api/images/convert
/api/images/crop
/api/images/rotate
/api/ocr/image
/api/ocr/pdf
Never trust:
FileName
Extension
Content-Type
from the client.
Validate:
Extension
+
Content-Type
+
Magic bytes/file signature
+
Maximum file size
+
Maximum page count
+
Processing limits
Also:
- Generate server-side storage keys.
- Sanitize original file names.
- Never use user-provided paths.
- Store uploads outside the web root.
- Never execute uploaded files.
- Apply request size limits.
- Use cancellation tokens.
- Clean temporary files.
- Scan uploads with an antivirus service when required by deployment policy.
Recommended lifecycle:
Uploaded
|
v
Stored
|
v
Queued
|
v
Processing
|
+---- Failed
|
v
Completed
|
v
Available for Download
|
v
Expired
|
v
Deleted
Files should have an expiration policy.
Example:
Anonymous files:
24 hours
Authenticated temporary files:
configurable retention
User-owned permanent files:
until user deletes them
Retention should be configurable.
Support both.
Upload
↓
Process
↓
Download
↓
Automatic expiration
No permanent file history is required.
Upload
↓
Process
↓
File History
↓
Download
↓
Manage Files
User ownership must be enforced server-side.
Every operation should be represented consistently.
Example:
public enum FileOperationType
{
MergePdf,
SplitPdf,
CompressPdf,
RotatePdf,
WatermarkPdf,
ProtectPdf,
OfficeToPdf,
PdfToWord,
PdfToImage,
ImageToPdf,
ResizeImage,
CompressImage,
Ocr
}This allows:
ProcessingJob
|
+-- OperationType
+-- Status
+-- Progress
+-- InputFiles
+-- OutputFiles
+-- Error
Example:
ProcessingJobs
--------------------------------
Id
UserId
OperationType
Status
Progress
StartedAt
CompletedAt
ErrorMessage
CreatedAt
Input/output relation:
ProcessingJob
|
+-- Input File 1
+-- Input File 2
|
+-- Output File
This allows merge/split/conversion operations to use different numbers of input and output files.
For production, separate API and workers.
Load Balancer
|
+---------+---------+
| |
v v
API 1 API 2
| |
+---------+---------+
|
v
Queue
|
+--------+--------+
| | |
v v v
Worker1 Worker2 Worker3
| | |
+--------+--------+
|
v
Object Storage
This allows the processing capacity to scale independently from the API.
Development:
Docker Compose
│
├── Angular
├── ASP.NET API
├── Worker
├── SQL Server
├── Redis
├── MinIO
├── Hangfire Dashboard
└── Nginx
Production:
Nginx / Cloud Load Balancer
|
v
ASP.NET API
|
v
Queue
|
v
Workers
|
+-----+------+
| |
v v
SQL Server Blob/S3
|
v
Redis
Because the backend is ASP.NET, Azure is a strong production option:
Angular
↓
Azure Static Web Apps / CDN
↓
ASP.NET Core API
↓
Azure App Service / Container Apps / AKS
↓
Queue / Hangfire
↓
Worker containers
↓
Azure Blob Storage
↓
Azure SQL Database
↓
Azure Cache for Redis
Alternative AWS:
CloudFront
↓
S3 / Angular
↓
ECS / EKS
↓
SQS
↓
Worker
↓
S3
↓
RDS SQL Server
↓
ElastiCache
Start with a simpler managed deployment and move to Kubernetes only when actual scale requires it.
Local:
Angular
ASP.NET Core
SQL Server
Redis
MinIO
Hangfire
LibreOffice
Ghostscript
Tesseract
Docker Compose
Example:
docker compose up -d
The application should be able to run without cloud credentials during local development.
Do not blindly install every package. Add packages only where the feature requires them.
Microsoft.EntityFrameworkCore
Microsoft.EntityFrameworkCore.SqlServer
Microsoft.EntityFrameworkCore.Design
FluentValidation
FluentValidation.DependencyInjectionExtensions
PdfPig
QuestPDF
Ghostscript and LibreOffice are external executables rather than normal NuGet-only dependencies.
Hangfire.Core
Hangfire.SqlServer
StackExchange.Redis
or the ASP.NET distributed-cache integration where appropriate.
SixLabors.ImageSharp
Magick.NET
Use an appropriate Tesseract .NET wrapper and keep the native OCR runtime available inside the worker container.
Serilog.AspNetCore
Test:
Application services
Validators
Business rules
File naming
Job state transitions
Metadata processing
Example:
MergePdfServiceTests
CompressPdfServiceTests
FileValidationServiceTests
ProcessingJobServiceTests
Test:
API
SQL Server
Storage
Redis
Background jobs
Example:
POST /api/pdf/merge
|
v
Application
|
v
Infrastructure
|
v
Output file
Do not make unit tests depend on real cloud storage.
Use structured logging.
Recommended:
Serilog
Every processing job should have:
CorrelationId
JobId
UserId
Operation
InputSize
OutputSize
Duration
Status
Error
Example:
JobId=abc123
Operation=CompressPdf
Status=Completed
DurationMs=8420
InputSize=18MB
OutputSize=4MB
Production monitoring should include:
API latency
API error rate
Queue length
Worker utilization
Processing duration
Storage failures
Database failures
Failed jobs
Disk/temp usage
File Upload
File Download
File Delete
File Validation
Storage abstraction
SQL metadata
Job model
Merge
Split
Extract Pages
Remove Pages
Reorder Pages
Rotate
Compress
Watermark
Page Numbers
Protect
Unlock with valid credentials
JPG → PDF
PDF → JPG
Resize
Compress
Crop
Rotate
Convert
Word → PDF
Excel → PDF
PowerPoint → PDF
Image → Text
PDF → Text
Scanned PDF → Searchable PDF
PDF → Word
PDF → Excel
PDF → PowerPoint
PDF/A
Repair
Redaction
Compare
Forms
Digital signatures
Controllers should not contain processing logic.
Application should not reference third-party processing libraries directly.
Infrastructure owns external tools.
Large files should use streaming wherever practical.
Actual files belong in object storage, not normal SQL rows.
Metadata belongs in SQL Server.
Redis is not file storage.
Long-running operations belong in background workers.
API and Worker should be independently scalable.
Every processing operation should have a consistent Job model.
All temporary files require an expiration/cleanup strategy.
External executables must run with controlled arguments and restricted permissions.
┌──────────────────────────────────────────────────────────────┐
│ ANGULAR UI │
│ │
│ PDF Tools | Conversion | Images | OCR | Account | History │
└─────────────────────────────┬────────────────────────────────┘
│ HTTPS
▼
┌──────────────────────────────────────────────────────────────┐
│ ASP.NET CORE API │
│ │
│ Controllers | Auth | Validation | Middleware | Swagger │
└─────────────────────────────┬────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────────┐
│ APPLICATION │
│ │
│ PDF | Conversion | Images | OCR | Files | Jobs | Business │
│ Rules + Interfaces │
└──────────────┬─────────────────────────────┬─────────────────┘
│ │
▼ ▼
┌──────────────────────────┐ ┌────────────────────────────┐
│ INFRASTRUCTURE │ │ BACKGROUND │
│ │ │ WORKERS │
│ EF Core │ │ │
│ Blob/S3/MinIO │ │ Hangfire │
│ PdfPig │ │ PDF Processing │
│ QuestPDF │ │ Office Conversion │
│ Ghostscript │ │ Image Processing │
│ LibreOffice │ │ OCR │
│ ImageSharp │ │ │
│ Magick.NET │ └──────────────┬─────────────┘
│ Tesseract │ │
│ Redis │ │
│ Serilog │ │
└────────────┬─────────────┘ │
│ │
▼ ▼
┌──────────────────────┐ ┌─────────────────────────┐
│ SQL SERVER │ │ OBJECT STORAGE │
│ │ │ │
│ Metadata │ │ Azure Blob / S3 │
│ Users │ │ MinIO (local) │
│ Jobs │ │ │
│ File records │ │ Actual files │
│ Audit │ │ │
└──────────────────────┘ └─────────────────────────┘
The finished project should demonstrate these industry-level concepts:
Clean Architecture
Dependency Injection
SOLID
Generic abstractions
ASP.NET Core Web API
Angular
EF Core
SQL Server
Object Storage
Cloud Architecture
File Streaming
PDF Processing
Office Conversion
Image Processing
OCR
Background Jobs
Redis
Distributed Processing
Docker
Docker Compose
Logging
Monitoring
Authentication
Authorization
Testing
CI/CD
Cloud Deployment
The key design decision is to make the system engine-independent:
Application
|
+-- IPdfProcessor
|
+-- IFileStorageService
|
+-- IOfficeConverter
|
+-- IImageProcessor
|
+-- IOcrService
|
+-- IJobService
|
v
Infrastructure
|
+---------+----------+
| | |
PdfPig Ghostscript LibreOffice
QuestPDF ImageSharp Tesseract
Azure Blob MinIO/S3
This makes the project easier to test, replace, scale, containerize, and deploy to cloud infrastructure.