The NLP Knowledge Discovery Platform is an end-to-end research prototype for scientific document exploration using Natural Language Processing (NLP), Knowledge Graphs, Information Retrieval, and Semantic Search techniques.
The system processes scientific publications from the arXiv dataset, extracts structured knowledge from unstructured text, constructs a knowledge graph representing relationships between documents, authors, topics, and entities, and enables both lexical and semantic retrieval of scientific information.
The project was developed as part of a Master's degree in Artificial Intelligence Engineering and implements the concepts proposed in:
Building a Knowledge Graph for Scientific Document Exploration using NLP and Semantic Retrieval
The primary objectives of this project are:
- Extract entities, concepts, keywords, and relations from scientific publications
- Construct a knowledge graph from extracted information
- Enable graph-based exploration of scientific literature
- Implement multiple retrieval approaches for document discovery
- Compare lexical, semantic, and graph-enhanced retrieval methods
- Evaluate retrieval quality using standard Information Retrieval metrics
- Demonstrate how structured knowledge improves scientific document exploration
Raw Scientific Documents
│
▼
Text Preprocessing
│
▼
Information Extraction
├── Named Entities
├── Keywords
└── Relations
│
▼
Knowledge Graph Construction
│
▼
Knowledge Graph
│
┌──────┼──────┐
▼ ▼ ▼
BM25 Semantic KG-Retrieval
│
▼
Evaluation
- Loading scientific documents from CSV, JSON, and JSONL formats
- Support for arXiv metadata datasets
- Dataset filtering and sampling
- Text cleaning and normalization
- Tokenization and lemmatization
- Stopword removal
- Named Entity Recognition (NER)
- TF-IDF keyword extraction
- TextRank keyword extraction
- Co-occurrence relation extraction
- Relation aggregation across documents
- Paper nodes
- Author nodes
- Entity nodes
- Topic nodes
- Relation edges
- Graph pruning and filtering
- GraphML export
- JSON export
- Interactive graph visualization
- Traditional lexical retrieval
- Ranking based on term frequency and inverse document frequency
- Sentence Transformer embeddings
- Embedding caching
- Vector similarity search
- Graph-based reranking
- Entity-aware retrieval
- Semantic graph expansion
- Latent Dirichlet Allocation (LDA)
- BERTopic integration
- Retrieval-Augmented Generation (RAG)
- Graph Question Answering
- Embedding Bias Analysis
- Debiasing Utilities
- Graph Neural Network Utilities
- Semantic Network Construction
- Streamlit dashboard
- Interactive graph exploration
- Topic visualization
- Search interface
- Network analysis
config.yaml
run_all.py
create_queries.py
src/
├── preprocessing/
├── extraction/
├── knowledge_graph/
├── retrieval/
├── experiments/
├── evaluation/
├── topic_modeling/
├── embeddings/
├── visualization/
├── llm_kg/
├── gnn/
└── bias/
tests/
data/
├── raw/
├── processed/
└── graphs/
reports/
├── figures/
└── tables/
python -m venv .venvWindows:
.venv\Scripts\activateLinux / macOS:
source .venv/bin/activatepip install --upgrade pip
pip install -r requirements.txtpython -m spacy download en_core_web_smpython run_all.py --full --skip-embeddingspython run_all.py --fullpython -m src.pipeline.run_pipeline --skip-embeddingspython run_all.py --input data/raw/arxiv-metadata-oai-snapshot.json --full --skip-embeddingsstreamlit run src/visualization/dashboard.pydata/processed/
├── processed_documents.csv
├── entities.csv
├── keywords.csv
├── relations.csv
└── graph_summary.json
data/graphs/
├── knowledge_graph.graphml
├── knowledge_graph.json
└── semantic_network.graphml
reports/tables/
├── bm25_results.csv
├── bm25_metrics.csv
├── semantic_results.csv
├── semantic_metrics.csv
├── kg_results.csv
└── kg_metrics.csv
reports/figures/
└── knowledge_graph.html
The project evaluates retrieval quality using:
- Precision@K
- Recall@K
- Mean Reciprocal Rank (MRR)
Experiments are performed for:
- BM25 Baseline Retrieval
- Semantic Retrieval
- Knowledge Graph Enhanced Retrieval
The repository includes automated tests covering:
- Data preprocessing
- Information extraction
- Retrieval systems
- Knowledge graph functionality
- Evaluation metrics
- Semantic retrieval
- Graph-enhanced retrieval
- Topic modeling helpers
- Additional extension modules
Run the tests:
pytest -v58 passed
- Python
- spaCy
- pandas
- NumPy
- scikit-learn
- NetworkX
- Sentence Transformers
- Streamlit
- PyTorch
- Gensim
The NLP Knowledge Discovery Platform successfully demonstrates how Natural Language Processing, Knowledge Graphs, and Semantic Retrieval can be combined to support scientific document exploration. The resulting system enables extraction of structured knowledge from scientific publications, graph-based representation of research concepts, and enhanced retrieval capabilities beyond traditional keyword search.
The platform provides a complete and reproducible workflow from raw scientific documents to searchable knowledge graphs and retrieval evaluation experiments.