Tech Lead · Applied Scientist · ML Researcher
I build research-driven machine-learning systems, from foundation-model architecture and rigorous evaluation to large-scale production deployment.
My work spans foundation models, representation learning, benchmarking, NLP, ranking, and genomic ML. I am the first author of two ICML 2026 main-track papers and currently lead conversion optimization at SberAds.
- Research Scientist at Moscow Independent Research Institute of AI (MIRAI)
- Tech Lead at SberAds
- Two-time winner of the International Olympiad on AI and Data Analysis (AIDAO)
- Main interests: foundation models, representation learning, model evaluation, ranking, dialogue systems, and genomic ML
- Email: a.ledn2026@gmail.com
- LinkedIn: https://www.linkedin.com/in/daria-ledneva
- Telegram: https://t.me/darlednik
GENEB: Why Genomic Models Are Hard to Compare
Accepted to ICML 2026, main track
GENEB is a diagnostic benchmark for genomic foundation models. It evaluates frozen DNA sequence representations across 100 classification tasks grouped into 13 functional categories, under full-data, 10-shot, and 1-shot regimes.
The benchmark shows that aggregate leaderboards can be misleading: model rankings vary strongly across biological task categories, and architecture or pretraining alignment can matter more than parameter count.
- Paper: https://arxiv.org/abs/2606.04525
- Code: https://github.com/darlednik/geneb
- Leaderboard: https://huggingface.co/spaces/darlednik/geneb-leaderboard
- Dataset: https://huggingface.co/datasets/darlednik/geneb-tasks
LDARNet: DNA Adaptive Representation Network with Learnable Tokenization for Genomic Modeling
Accepted to ICML 2026, main track
LDARNet is a 110M-parameter hierarchical genomic foundation model with learnable tokenization. It adapts dynamic chunking to masked language modeling through bidirectional routing and bidirectional dechunking.
Across downstream genomic tasks, LDARNet shows that learned token boundaries can improve compact genomic encoders, especially on long-range regulatory and epigenetic tasks.
- Paper: https://arxiv.org/abs/2606.04552
- Code: https://github.com/darlednik/ICML-LDARNet
- Models: https://huggingface.co/collections/darlednik/ldarnet
-
LDARNet: DNA Adaptive Representation Network with Learnable Tokenization for Genomic Modeling
ICML 2026, main track
https://arxiv.org/abs/2606.04552 -
GENEB: Why Genomic Models Are Hard to Compare
ICML 2026, main track
https://arxiv.org/abs/2606.04525 -
Joint Distillation and Contrastive Learning for Embeddings in Task-Oriented Dialogue Systems
Doklady Mathematics, 2025
https://link.springer.com/article/10.1134/S1064562425700243 -
Reimagining Intent Prediction: Insights from Graph-Based Dialogue Modeling and Sentence Encoders
LREC-COLING 2024
https://aclanthology.org/2024.lrec-main.1208/ -
Graph Models for Contextual Intention Prediction in Dialog Systems
Doklady Mathematics, 2024
https://link.springer.com/article/10.1134/S106456242370117X
- Genomic foundation models
- Representation learning
- Benchmark design and evaluation
- Dialogue systems and task-oriented dialogue
- NLP and LLM-based systems
- Applied ML in production
Languages and ML: Python, PyTorch, Hugging Face Transformers, Sentence Transformers, LoRA / PEFT
Data and production: PySpark, SQL, Airflow, PostgreSQL, Docker, Git, CI/CD
Research workflows: benchmarking, model evaluation, probing, reproducibility, dataset construction
- Two-time winner, International Olympiad on AI and Data Analysis (AIDAO), 2024 & 2025
- Bronze medalist, “I Am a Professional” Olympiad in Artificial Intelligence, 2024
- Invited speaker, Aha 2025
- Featured speaker, Huawei Workshop on ML/AI/NLP for search-engine performance, 2022
- Winner, YaProfi Hackathon at the MIPT Educational Forum, 2023
- Winner, “Digital Methods in Energy” Hackathon, STC Gazprom Neft


