Skip to content

About

🗃 Classification of toxic and neutral messages on russian

Resources

Stars

3 stars

Watchers

1 watching

Forks

Repository files navigation

Russian toxic messages classification

Lightweight multi-label classifier for Russian text toxicity. Linear model on word- and char-level TF-IDF — no transformers, no GPU, ms-level inference, small model footprint.

gradio

Quick example

from api import API
api = API()

api.check('ты п###р')
# scores:    toxic=1.000  profanity=0.03  insult=0.73  identity_attack=1.00
# verdicts:  toxic, identity_attack

api.check('дeбил, что ты несёшь')          # обфускация ловится char n-grams
# scores:    toxic=0.999  profanity=0.07  insult=0.97  identity_attack=0.00
# verdicts:  toxic, insult

api.check('бля, как красиво')              # мат-эмоция, не атака
# scores:    toxic=0.31  profanity=0.96  insult=0.00  identity_attack=0.01
# verdicts:  profanity (only)

api.check('х###ы опять что-то придумали')
# scores:    toxic=0.98  profanity=0.02  insult=0.00  identity_attack=1.00
# verdicts:  toxic, identity_attack

api.check('спасибо большое за помощь!')
# scores:    toxic=0.02  profanity=0.00  insult=0.00  identity_attack=0.01
# verdicts:  (none — neutral)

Full response shape:

{
    'scores':     {'toxic': 1.0,  'profanity': 0.03, 'insult': 0.73, 'identity_attack': 1.0},
    'verdicts':   {'toxic': True, 'profanity': False, 'insult': False, 'identity_attack': True},
    'thresholds': {'toxic': 0.31, 'profanity': 0.14, 'insult': 0.76, 'identity_attack': 0.20},
    'work_time':  '0.04s',
}

Categories

Four independent heads. A message can fire multiple categories at once. Each category has its own P/R-tuned threshold (default precision target 0.85).

Category Meaning Fires on
toxic Общая токсичность сообщения — модель учится на оригинальной бинарной разметке. Покрывает кейсы без явных лексических маркеров (сарказм, агрессивный тон). «иди отсюда, надоел», «опять ты со своим бредом»
profanity Содержит русский мат — независимо от того, направлен ли он на кого-то. «бля, как красиво», «дохуя народу»
insult Направленное оскорбление человеку (insult-слово + 2-е лицо или императив). «ты дебил», «иди сдохни», «вы тупые»
identity_attack Атака на группу — нац./гендер/ориентация/идентичность. «хохлы», «пидорас», «ватники», «жиды»

Запросы вроде 'бля, как красиво' правильно отлетают как profanity-only, без toxic/insult — это и есть смысл multi-label.

Architecture

text
 └─ normalize (homoglyph fold, repeat collapse, junk strip)
 └─ FeatureUnion
      ├─ word TF-IDF  (lemmatized via pymorphy3, ngrams 1–2)
      └─ char_wb TF-IDF  (ngrams 3–5)
 └─ PerHeadLogReg  (one LogisticRegression per label,
                    rare classes get class_weight='balanced')
 └─ per-head threshold (P/R-tuned, default target precision 0.85)

The whole inference pipeline is a single sklearn.Pipeline saved with dill. Defined in model_utils.py.

Quality (held-out 10% test, ~14k labelled comments)

Head F1 Precision Recall Notes
toxic ~0.88 ~0.85 ~0.92 Direct upgrade of the original binary baseline
profanity ~0.90 ~0.85 ~0.95 Strongest head — char n-grams nail explicit lexicon
insult ~0.85* — — Rare class (~1.8%), class_weight='balanced' enabled
identity_attack ~0.80* — — Rare class (~4.5%), class_weight='balanced' enabled

* macro F1 across 'no'/'yes' classes; positive-class F1 varies per run.

Repo layout

labeled.csv                       # original binary-labeled corpus (~14k comments)
labeled_multi.csv                 # multi-label labels produced by make_multilabel.py
make_multilabel.py                # rule-based weak-supervision relabeller
model_utils.py                    # shared definitions (normalize, tokenize, PerHeadLogReg)
tfidf_logreg_multilabel.ipynb     # current production model: training + threshold tuning
tfidf_logreg_classifier.ipynb     # earlier binary model + data-cleaning analysis (kept for reference)
api.py                            # Python API around the saved pipeline
bot.py                            # aiogram Telegram bot
server.py                         # Gradio web UI (port 7860)
pipeline_multilabel.pkl           # trained pipeline (produced by the notebook)
thresholds_multilabel.pkl         # per-head thresholds (produced by the notebook)

Running

For Python usage see Quick example above.

Gradio web UI

pip install -r requirements.txt
python server.py
# open http://localhost:7860

Telegram bot

export BOT_TOKEN=...
python bot.py

Docker

docker build -t ru-toxic .
docker run -p 7860:7860 ru-toxic

Retraining

The model is trained from scratch in two steps.

1. Generate multi-label data from the original binary labeled.csv:

python make_multilabel.py
# -> labeled_multi.csv

This is a rule-based weak-supervision pass — lexicons + regex over normalized text. Quality on a manual audit ≈ 80–85% per category; sufficient to bootstrap ML training. Replace with cleaner labels (LLM relabel, manual annotation, or a real multi-label dataset like Jigsaw RU) for production-grade training.

2. Train the model: open tfidf_logreg_multilabel.ipynb and run all cells. This writes pipeline_multilabel.pkl and thresholds_multilabel.pkl consumed by api.py.

Design notes

  • No transformers by design. The model targets pattern detection (n-grams, lemmas, lexicons), not contextual understanding. This is a feature: small (~1–10 MB), fast (ms per call), interpretable, runs anywhere.
  • Char n-grams + Cyrillic↔Latin homoglyph fold make the model robust against typical obfuscations: пuдор, д*бил, КЛАССССС.
  • Per-head class_weight — class_weight='balanced' is applied only to rare-class heads (insult 1.8%, identity_attack 4.5%) to avoid trivially- zero predictions, while toxic and profanity train without it for cleaner probability calibration. Lives in HEAD_CONFIGS inside the notebook.
  • Per-head thresholds. Each head has its own P/R-tuned threshold — there's no good reason to use 0.5 for a 33%-positive class and a 1.8%-positive class simultaneously. Thresholds are stored in thresholds_multilabel.pkl alongside the pipeline.

Limitations

  • Contextual toxicity without lexical markers (sarcasm, sneer with no obscenity, hostile rhetorical questions) is the architectural ceiling of pattern-based models. ~60% of comments labeled toxic=1 in the source dataset have no explicit toxic markers — the model still catches most of them via learned correlations, but with FPs in adjacent neutral contexts.
  • Greeting phrases like "Привет, как дела?" can fire toxic because similar phrases appear in sarcastic/aggressive contexts in the training data. This is a property of the dataset, not the architecture; LLM relabelling or manual cleaning would resolve it.
  • The threat rule-derived label has only ~24 positives — it is excluded from ML training; if a threat detector is needed, use the rule-based detector from make_multilabel.py directly.

License

MIT — see LICENSE.

About

🗃 Classification of toxic and neutral messages on russian

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Used by

Contributors

Languages