Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
44 changes: 41 additions & 3 deletions .github/copilot-instructions.md
Original file line number Diff line number Diff line change
@@ -1,3 +1,44 @@
# Copilot instructions for Scrape-UAlg-Courses

Goal: a complete scraper + REST API for UAlg courses. Source pages are scraped into a normalized SQLite DB; the FastAPI service reads from that DB and serves a simple web UI.

## Architecture and data flow
- Scraper (basic) in `src/scraper.py` (class `UAlgScraper`): fetches HTML and parses simple "course" blocks. Used mainly in unit tests/examples.
- Scraper (full) in `src/scrape_ualg.py` (class `UAlgCourseScraper`):
- Initializes DB from `schema.sql` via `init_db()`.
- Crawls listing pages starting at `START_URL`, extracts course links, parses each course page into fields: title, code, level, school, language, modules, areas, documents.
- Persists to SQLite (`ualg_courses.db` by default), also downloads documents to `data/docs/`.
- Respects retries (`Config.max_retries`) and rate limiting (sleep ~1.5s between courses).
- API in `src/api.py` (FastAPI): reads from the same SQLite DB (`DB_PATH = "ualg_courses.db"`) and serves endpoints (`/courses`, `/courses/{id}`, `/levels`, `/schools`, `/areas`, `/stats`, `/api`, `/`). The `/` route serves `templates/index.html`.
- Frontend in `templates/index.html`: fetches API endpoints and renders dashboards with Chart.js.

## Dev workflows (Windows PowerShell friendly)
- Install deps: `make install`; dev deps: `make install-dev`.
- Initialize DB: `make init-db`; load demo data: `make demo` (runs `scripts/populate_demo_data.py`).
- Run scraper: `make run-scraper` (equivalent to `python -m src.scrape_ualg`).
- Run API: `make run-api` (uvicorn `src.api:app` on port 8000).
- Tests: `make test` (pytest + coverage). Lint/format: `make lint`, `make format`.

## Conventions and patterns
- Config is centralized in `src/config.py` (`Config`): defaults to `https://www.ualg.pt`, timeout=30, max_retries=3, UA set; override with env `UALG_BASE_URL` or via constructor.
- Logging: both scrapers log progress; keep user-agent and backoff semantics; do not remove the ~1.5s delay in the full scraper.
- DB access in API: use `get_db_connection()` with `row_factory = sqlite3.Row` and helper `dict_from_row`.
- Filtering: `/courses` builds SQL dynamically with optional joins on `levels`, `schools`, and `areas`; keep `limit` (1..500) and `offset` semantics.
- Schema is authoritative: see `schema.sql`. M2M tables: `course_area`, `module_course`; uniqueness and indices are already defined.
- Paths: code locates `schema.sql` and `templates/index.html` via `..` from `src/`; keep relative paths when adding files.

## When extending
- API: mirror existing patterns (Pydantic response models at top; small helpers; parameterized SQL; close connections). Add tags and docs to endpoints.
- Scraper: prefer adding selectors heuristically (see `parse_course_page`), and persist via existing upsert helpers; keep downloads in `data/docs/`.
- Tests: look at `tests/test_api.py`, `tests/test_scraper.py`, `tests/test_scrape_ualg.py` for expected behavior and monkeypatch patterns (e.g., override `DB_PATH` in API tests).

## Quick examples
- Add an API filter: extend `list_courses` by appending the join + where + param; preserve `limit/offset` bounds.
- Add a new field to courses: update `schema.sql` + insert in `save_course` + select in API queries + include in models.

Notes
- Default DB filename is `ualg_courses.db` in repo root. Populate it via scraper or `scripts/populate_demo_data.py` before using the UI.
- CI: `.github/workflows/python-app.yml` runs tests/lint on PRs; match the Makefile targets.
## Purpose

This file gives short, actionable guidance for AI coding agents working in this repository so they can be productive immediately.
Expand Down Expand Up @@ -41,6 +82,3 @@ This file gives short, actionable guidance for AI coding agents working in this

## When to ask the maintainer
- If the target site structure (selectors, endpoints) is unknown, ask for a sample HTML page or the desired output schema.

---
If anything here is unclear or you'd like more examples (e.g., a test change or a small scraper tweak), tell me which section to expand.
51 changes: 51 additions & 0 deletions .github/workflows/copilot-setup-steps.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
name: "Copilot Setup Steps"

# Automatically run the setup steps when they are changed to allow for easy validation,
# and allow manual testing through the repository's "Actions" tab
on:
workflow_dispatch:
push:
paths:
- .github/workflows/copilot-setup-steps.yml
pull_request:
paths:
- .github/workflows/copilot-setup-steps.yml

jobs:
# The job MUST be called copilot-setup-steps or it will not be picked up by Copilot.
copilot-setup-steps:
runs-on: ubuntu-latest

# Set the permissions to the lowest permissions possible needed for your steps.
# Copilot will be given its own token for its operations.
permissions:
# If you want to clone the repository as part of your setup steps, for example to install dependencies,
# you'll need the `contents: read` permission. If you don't clone the repository in your setup steps,
# Copilot will do this for you automatically after the steps complete.
contents: read

# You can define any steps you want, and they will run before the agent starts.
# If you do not check out your code, Copilot will do this for you.
steps:
- name: Checkout code
uses: actions/checkout@v4

- name: Set up Python 3.10
uses: actions/setup-python@v5
with:
python-version: "3.10"
cache: "pip"

- name: Install Python dependencies
run: |
python -m pip install --upgrade pip
pip install -r requirements.txt
pip install -r requirements-dev.txt

- name: Initialize database
run: |
python -c "from src.scrape_ualg import UAlgCourseScraper; s = UAlgCourseScraper(); s.init_db()"

- name: Verify installation
run: |
python -c "import requests; import bs4; import fastapi; import pytest; print('All dependencies installed successfully')"
6 changes: 5 additions & 1 deletion Makefile
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
.PHONY: help install install-dev test lint format clean run run-scraper run-api init-db
.PHONY: help install install-dev test lint format clean run run-scraper run-api init-db demo

help:
@echo "Available targets:"
Expand All @@ -12,6 +12,7 @@ help:
@echo " run-scraper - Run the scraper"
@echo " run-api - Run the API server"
@echo " init-db - Initialize the database"
@echo " demo - Populate database with demo data"

install:
pip install -r requirements.txt
Expand Down Expand Up @@ -49,3 +50,6 @@ run-api:

init-db:
python -c "from src.scrape_ualg import UAlgCourseScraper; s = UAlgCourseScraper(); s.init_db()"

demo:
python scripts/populate_demo_data.py
57 changes: 51 additions & 6 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,10 +7,14 @@ Sistema completo de scraping e API REST para cursos da Universidade do Algarve (
- 🚀 **Scraper Completo**: Extrai informações detalhadas de cursos da UAlg
- 💾 **Banco de Dados SQLite**: Armazena dados em schema normalizado
- 🌐 **API REST com FastAPI**: Endpoints para consultar cursos, níveis, escolas e áreas
- 🎨 **Interface Web Moderna**: Frontend responsivo com visualizações interativas
- 📊 **Estatísticas e Gráficos**: Dashboard com Chart.js para visualização de dados
- 📄 **Download de Documentos**: Faz download automático de PDFs e documentos
- 🎓 **Extração de Módulos/UCs**: Captura unidades curriculares com ECTS, ano, semestre
- 🏷️ **Áreas de Conhecimento**: Organiza cursos por áreas temáticas
- 🔧 **Configurável**: Timeouts, retries, user agents personalizáveis
- 📦 **Modular**: Código organizado e separado em módulos
- ✅ **Testes Completos**: Cobertura abrangente de testes
- ✅ **Testes Completos**: Cobertura abrangente de testes (36 testes)
- 🔄 **Retry Automático**: Mecanismo de retry para requisições falhadas
- 📝 **Logging Detalhado**: Logs para debugging e monitoramento
- ⚡ **Rate Limiting**: Respeita limites do servidor com delays entre requisições
Expand Down Expand Up @@ -81,6 +85,23 @@ scraper.init_db()

## Uso

### Quick Start com Dados de Demonstração

Para testar rapidamente o sistema com dados de exemplo:

```bash
# 1. Inicializar banco de dados
make init-db

# 2. Popular com dados de demonstração
make demo

# 3. Iniciar servidor API
make run-api

# 4. Acessar interface web em http://localhost:8000/
```

### Scraper Básico (Original)

```python
Expand Down Expand Up @@ -138,10 +159,21 @@ uvicorn src.api:app --reload --host 0.0.0.0 --port 8000

A API estará disponível em `http://localhost:8000`

#### Interface Web

Acesse `http://localhost:8000/` para visualizar a interface web moderna com:
- 📊 Dashboard com estatísticas gerais
- 🔍 Filtros interativos por nível, escola e área
- 📈 Gráficos de distribuição de cursos
- 🎴 Cards de cursos com informações detalhadas
- 📱 Design responsivo

#### Documentação Interativa da API

- Interface Web: `http://localhost:8000/`
- Swagger UI: `http://localhost:8000/docs`
- ReDoc: `http://localhost:8000/redoc`
- API Info (JSON): `http://localhost:8000/api`

#### Exemplos de Uso da API

Expand All @@ -165,6 +197,11 @@ curl "http://localhost:8000/courses?level=Licenciatura"
curl "http://localhost:8000/courses?school=Faculdade%20de%20Ciências"
```

**Filtrar cursos por área:**
```bash
curl "http://localhost:8000/courses?area=Tecnologias%20da%20Informação"
```

**Obter estatísticas:**
```bash
curl http://localhost:8000/stats
Expand Down Expand Up @@ -241,6 +278,8 @@ Scrape-UAlg-Courses/
│ ├── scraper.py # Scraper básico (original)
│ ├── scrape_ualg.py # Scraper completo com BD
│ └── api.py # API REST FastAPI
├── templates/
│ └── index.html # Interface web moderna
├── tests/
│ ├── __init__.py
│ ├── test_scraper.py # Testes do scraper básico
Expand Down Expand Up @@ -318,9 +357,10 @@ O scraper pode ser configurado usando a classe `Config` ou variáveis de ambient

| Método | Endpoint | Descrição |
|--------|----------|-----------|
| GET | `/` | Informações da API |
| GET | `/` | Interface web moderna (HTML) |
| GET | `/api` | Informações da API (JSON) |
| GET | `/courses` | Listar cursos (com filtros opcionais) |
| GET | `/courses/{id}` | Obter detalhes de um curso |
| GET | `/courses/{id}` | Obter detalhes de um curso com módulos, áreas e documentos |
| GET | `/levels` | Listar níveis de curso |
| GET | `/schools` | Listar escolas/faculdades |
| GET | `/areas` | Listar áreas de conhecimento |
Expand All @@ -331,7 +371,7 @@ O scraper pode ser configurado usando a classe `Config` ou variáveis de ambient
- `level`: Filtrar por nível (ex: "Licenciatura", "Mestrado")
- `school`: Filtrar por escola
- `area`: Filtrar por área de conhecimento
- `limit`: Número máximo de resultados (padrão: 100)
- `limit`: Número máximo de resultados (padrão: 100, máx: 500)
- `offset`: Offset para paginação (padrão: 0)

## Boas Práticas e Considerações
Expand Down Expand Up @@ -374,7 +414,7 @@ Se encontrar problemas ou tiver dúvidas:

## Agradecimentos

- Construído com [Requests](https://requests.readthedocs.io/), [Beautiful Soup](https://www.crummy.com/software/BeautifulSoup/) e [FastAPI](https://fastapi.tiangolo.com/)
- Construído com [Requests](https://requests.readthedocs.io/), [Beautiful Soup](https://www.crummy.com/software/BeautifulSoup/), [FastAPI](https://fastapi.tiangolo.com/) e [Chart.js](https://www.chartjs.org/)
- Inspirado pela necessidade de fácil acesso às informações de cursos da UAlg

## Roadmap
Expand All @@ -383,7 +423,12 @@ Se encontrar problemas ou tiver dúvidas:
- [x] Banco de dados SQLite normalizado
- [x] API REST completa
- [x] Download automático de documentos
- [x] Testes completos
- [x] Testes completos (36 testes)
- [x] Extração de módulos/UCs com ECTS, ano, semestre
- [x] Extração e organização por áreas de conhecimento
- [x] Interface web moderna e responsiva
- [x] Dashboard com estatísticas e gráficos
- [x] Filtros interativos por nível, escola e área
- [ ] Interface de linha de comando (CLI)
- [ ] Exportação de dados (JSON, CSV, Excel)
- [ ] Suporte para agendamento automático
Expand Down
Loading