Skip to content

Add Docker Compose architecture with 4 microservices and Hugo/FACODI export system - #6

Merged
marcelo-m7 merged 4 commits into
mainfrom
copilot/analyze-ualg-website
Oct 27, 2025
Merged

marcelo-m7 merged 4 commits into
mainfrom
copilot/analyze-ualg-website

Conversation

Copilot AI commented Oct 27, 2025 •

Copy link
Copy Markdown

Implements Docker-based production architecture with isolated services for API, scraping, and Hugo site generation, plus automated export pipeline for FACODI.pt integration.

Infrastructure

  • Docker Compose stack: 4 services (API, Scraper, FACODI/Hugo, shared SQLite volume)
  • Dockerfiles: Optimized builds for FastAPI, Python scraper, Hugo+nginx
  • Configuration: Environment-based setup with .env.example, rate limiting support (REQUEST_DELAY)
  • Networking: Health checks, service dependencies, bridge network

Export System

  • export_to_hugo.py: SQLite → Hugo markdown with YAML front matter, module organization, document copying
  • sync_with_facodi.sh: Automation helper with validation and statistics
  • Hugo layouts: Complete examples (list/single/summary templates, CSS, JS filters)

API Enhancements

New endpoints for FACODI integration:

GET /api/v1/courses/export          # Bulk JSON export with metadata
GET /api/v1/course/{id}/markdown    # Single course as markdown

Documentation (33,700 words)

  • docs/INTEGRATION.md: Architecture, data flow, setup workflow, troubleshooting
  • docs/HUGO_LAYOUTS_EXAMPLE.md: Complete template examples with styling
  • docs/QUICK_REFERENCE.md: Command cheat sheet, common tasks
  • docs/IMPLEMENTATION_SUMMARY.md: Technical overview
  • README.md: Docker setup, FACODI integration workflow, updated endpoints

Quick Start

cp .env.example .env
make docker-up                # Start API + FACODI
make docker-scraper          # Populate database
./scripts/sync_with_facodi.sh  # Export to Hugo

Services: API (8000), FACODI (3000), on-demand scraper (profile-based), shared SQLite.

Testing

All 46 tests passing. No breaking changes. Backwards compatible with local development.

Original prompt

Análise do site UAlg e Plano de Implementação

1) Investigação do site UAlg - URLs coletadas ✅

URLs Exemplares Identificadas:

Páginas índice/listagem:

Exemplos de páginas de curso (padrão /curso/{ID}):

Análise técnica:

  • ✅ robots.txt: Disponível em https://www.ualg.pt/robots.txt
  • ✅ Estrutura: URLs seguem padrão previsível /curso/{ID} com ID numérico
  • ✅ Conteúdo por página: plano curricular, objetivos, ECTS, regime, área temática, saídas profissionais, contactos, documentos
  • ⚠️ JavaScript: Páginas parecem ser server-side rendered (confirmação necessária ao testar scraper)

2) Arquitetura Docker Compose - 4 Serviços Isolados

Baseado nas suas especificações, vou estruturar o projeto com:

Serviço 1: API Python + SQLite

  • FastAPI servindo endpoints REST
  • SQLite como DB persistente (volume compartilhado)
  • Serve frontend simples de visualização
  • Porta: 8000

Serviço 2: Frontend Visualização Scraper

  • Interface simples HTML/JS para visualizar dados do scraper
  • Consome API Python
  • Integrado com o serviço 1 (servido pela FastAPI)

Serviço 3: Scraper Service

  • Serviço independente que roda scraper periodicamente
  • Conecta-se à API Python para atualizar dados
  • Pode ser acionado por cron/schedule ou manualmente
  • Compartilha volume do SQLite

Serviço 4: Frontend FACODI.pt (clone)

  • Site estático Hugo
  • Clone do repositório facodi.pt dentro do Scrape-UAlg-Courses
  • Consome API Python para dados dinâmicos
  • Build de produção servido por nginx
  • Porta: 3000

3) Estrutura de Diretórios Proposta

Scrape-UAlg-Courses/
├── docker-compose.yml
├── .env.example
├── Makefile (atualizado para Docker)
├── src/
│   ├── api.py (FastAPI)
│   ├── scraper.py
│   ├── scrape_ualg.py
│   ├── config.py
│   └── ...
├── templates/
│   └── index.html (frontend visualização)
├── data/
│   ├── ualg_courses.db (SQLite - volume persistente)
│   └── docs/ (documentos baixados)
├── facodi-clone/ (git clone do facodi.pt)
│   ├── content/
│   ├── layouts/
│   ├── static/
│   ├── config.toml
│   └── ...
├── docker/
│   ├── api.Dockerfile
│   ├── scraper.Dockerfile
│   ├── facodi.Dockerfile
│   └── nginx.conf (para servir Hugo build)
└── scripts/
    ├── export_to_hugo.py (novo - exporta DB → MD)
    ├── populate_demo_data.py
    └── sync_with_facodi.sh

4) Próximas Etapas Concretas (Priorizadas)

Fase 1: Preparação e Infraestrutura (Imediato - 1-2 dias)

1.1. Verificar robots.txt e permissões

curl https://www.ualg.pt/robots.txt
  • Confirmar paths permitidos para scraping
  • Documentar rate limits recomendados
  • ✅ Saída esperada: Documento com regras de scraping permitidas

1.2. Clonar facodi.pt para dentro do repositório

cd Scrape-UAlg-Courses
git clone https://github.com/Monynha-Softwares/facodi.pt.git facodi-clone
cd facodi-clone
git checkout master
  • ✅ Saída esperada: facodi-clone/ directory com código Hugo

1.3. Criar estrutura Docker

  • Criar docker-compose.yml com 4 serviços
  • Criar Dockerfiles em docker/
  • Configurar volumes compartilhados (SQLite, docs)
  • ✅ Saída esperada: docker-compose up sobe todos os serviços

1.4. Configurar variáveis de ambiente

cp .env.example .env

Variáveis necessárias:

UALG_BASE_URL=https://www.ualg.pt
DB_PATH=/data/ualg_courses.db
MAX_RETRIES=3
REQUEST_DELAY=0.5
API_PORT=8000
FACODI_PORT=3000

Fase 2: Desenvolvimento do Scraper e API (Curto prazo - 1 semana)

2.1. Ajustar scraper para URLs identificadas

  • Atualizar START_URL em scrape_ualg.py
  • Implementar parser para /oferta-formativa e /licenciaturas
  • Testar com as 4 URLs exemplares coletadas
  • ✅ Saída esperada: Scraper extrai dados das páginas de exemplo

2.2. Implementar rate limiting configurável

# em config.py
REQUEST_DELAY = float(os.getenv('REQUEST_DELAY', '0.5'))  # 0.5s = 2 req/sec
  • Adicionar delay entre requests
  • Implementar backoff exponencial (3 retries)
  • ✅ Saída esperada: Logs mostrando delays respeitados

2.3. Expandir schema.sql (se necessário)

  • Adicionar campos que faltam conforme páginas reais
  • Campos sugeridos: duration_years, ects_total, regime, candidature_info
  • ✅ Saída esperada: Schema atualizado e migrações documentadas

2.4. API: Adicionar endpoints para FACODI

# Novos endpoints em api.py
@app.get("/api/v1/courses/export")  # Exportar todos cursos para ...

</details>



<!-- START COPILOT CODING AGENT TIPS -->
---

✨ Let Copilot coding agent [set things up for you](https://github.com/Monynha-Softwares/Scrape-UAlg-Courses/issues/new?title=✨+Set+up+Copilot+instructions&body=Configure%20instructions%20for%20this%20repository%20as%20documented%20in%20%5BBest%20practices%20for%20Copilot%20coding%20agent%20in%20your%20repository%5D%28https://gh.io/copilot-coding-agent-tips%29%2E%0A%0A%3COnboard%20this%20repo%3E&assignees=copilot) — coding agent works faster and does higher quality work when set up for your repo.

Copilot AI and others added 3 commits October 27, 2025 12:03
Co-authored-by: marcelo-m7 <117441129+marcelo-m7@users.noreply.github.com>
…ples

Co-authored-by: marcelo-m7 <117441129+marcelo-m7@users.noreply.github.com>
Co-authored-by: marcelo-m7 <117441129+marcelo-m7@users.noreply.github.com>
Copilot AI changed the title [WIP] Analyze UAlg website and outline implementation plan Add Docker Compose architecture with 4 microservices and Hugo/FACODI export system Oct 27, 2025
Copilot AI requested a review from marcelo-m7 October 27, 2025 12:14
@marcelo-m7
marcelo-m7 marked this pull request as ready for review October 27, 2025 12:17
@marcelo-m7
marcelo-m7 merged commit 50bff1a into main Oct 27, 2025
2 checks passed
@marcelo-m7
marcelo-m7 deleted the copilot/analyze-ualg-website branch October 27, 2025 12:17

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/api.py
Comment on lines 20 to 21
# Configurações
DB_PATH = "ualg_courses.db"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Respect DB_PATH env so Docker services share database

The new Compose stack mounts ./data:/data and passes DB_PATH=/data/ualg_courses.db, but the API still hard‑codes DB_PATH = "ualg_courses.db". When the container starts, SQLite creates /app/ualg_courses.db instead of using the shared /data volume, so the API never sees data written by the scraper and the database is not persisted. Reading the path from the environment (or the Config class) would allow both services to operate on the same file.

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants