Pinned Loading
-
TimeSeriesForecasting
TimeSeriesForecasting PublicA production-grade benchmark framework designed to stress-test deep time-series models (DLinear, Linear_v2) against zero-parameter baselines. It simulates live telemetry loops by streaming sequence…
Python
-
-
DepthPruning-vLLM
DepthPruning-vLLM PublicDeployment-aware transformer compression pipeline using structured depth pruning, knowledge distillation, and vLLM serving. Benchmarks analyze latency, throughput, concurrency scaling, and KV-cache…
-
InferenceScaling-Qwen
InferenceScaling-Qwen PublicInvestigating how inference-time compute influences reasoning performance in large language models through systematic scaling of self-consistency decoding.
Python
-
Triton-FlashAttention
Triton-FlashAttention PublicGQA/MQA-extended FlashAttention-2 in Triton, with a KV-cache decode kernel — built on a public tutorial baseline, hardened for correctness (4 bugs found/fixed), and benchmarked against PyTorch SDPA…
Python
If the problem persists, check the GitHub status page or contact support.