Skip to content

Commit 72012de

Browse files
committed
Update docs with temperature and ResNet50 findings
- Temperature ablation: T=4 optimal, T=6/T=8 perform worse - ResNet50 conv1+KD: 77-81% gap recovery (vs 88-123% for ResNet18) - Bump experiment count to 153 - Update CLAUDE.md with lambda server workflow
1 parent c570cd5 commit 72012de

2 files changed

Lines changed: 180 additions & 37 deletions

File tree

‎CLAUDE.md‎

Lines changed: 10 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -14,7 +14,7 @@ Research project studying **BitNet b1.58** (1.58-bit ternary quantization) appli
1414

1515
## Project Structure
1616

17-
```
17+
```text
1818
experiments/
1919
├── train.py # Standard training script
2020
├── train_kd.py # Knowledge Distillation training
@@ -59,7 +59,7 @@ uv run python -m analysis.generate_figures
5959

6060
## Results Directory Structure
6161

62-
```
62+
```text
6363
results/raw/{dataset}/{model}/{version}_s{seed}/
6464
├── config.json # Training config
6565
├── results.json # Final metrics
@@ -69,15 +69,20 @@ results/raw/{dataset}/{model}/{version}_s{seed}/
6969

7070
## Server Workflow
7171

72-
Code lives locally, experiments run on remote server with GPUs.
72+
Code lives locally, experiments run on `lambda` server with GPUs.
7373

7474
```bash
75-
# On server: archive results (JSON only, not model weights)
75+
# On server (lambda): archive results (JSON only, not model weights)
7676
find results/raw -name "*.json" | tar -czvf results_json.tar.gz -T -
7777

7878
# Locally: download and extract
79-
scp server:~/code/.../results_json.tar.gz .
79+
scp lambda:/home/dcazzani/code/lab-strange-loop/bitnet/results_json.tar.gz .
8080
tar -xzvf results_json.tar.gz
81+
82+
# Regenerate analysis artifacts
83+
uv run python -m analysis.aggregate_results # reads raw/, saves processed/
84+
uv run python -m analysis.generate_tables # reads raw/, writes paper/tables/
85+
uv run python -m analysis.generate_figures # reads raw/, writes paper/figures/
8186
```
8287

8388
## Important Notes

‎PLAN.md‎

Lines changed: 170 additions & 32 deletions
Original file line numberDiff line numberDiff line change
@@ -1,8 +1,36 @@
11
# BitNet CNN Research Plan
22

3-
**Last Updated**: Feb 2, 2026
4-
**Primary Target**: WACV 2027 Round 2 (Sept 2026)
3+
**Last Updated**: Feb 5, 2026
4+
**Primary Target**: CVPR 2026 Workshop (~Apr 2026)
55
**Parallel Track**: NeurIPS 2026 Efficient ML Workshop (Aug 2026)
6+
**Backup**: WACV 2027 Round 2 (Sept 2026)
7+
8+
---
9+
10+
## Paper Status: 📝 FIRST DRAFT COMPLETE
11+
12+
**Title**: "When Augmentation Fails: Knowledge Distillation for Ternary CNNs"
13+
14+
**Paper file**: `paper/main.tex` (13 pages, builds successfully)
15+
16+
### Sections Completed
17+
18+
- ✅ Title and author info (Dario Cazzani, Cisco Systems Inc.)
19+
- ✅ Abstract
20+
- ✅ Introduction (framing the augmentation paradox)
21+
- ✅ Related Work (with TTQ, QKD, HAQ citations)
22+
- ✅ Method (BitNet b1.58 implementation details)
23+
- ✅ Section 4: "What Doesn't Work: The Augmentation Paradox"
24+
- ✅ Section 5: "What Works: A Practical Recipe" (conv1 + KD)
25+
- ✅ Discussion
26+
- ✅ Conclusion
27+
- ✅ Reproducibility Appendix (code → results → paper pipeline)
28+
29+
### Still Needed
30+
31+
- 📝 Complete figures (Fig 4: Recipe comparison bar chart)
32+
- 📝 Polish and proofread
33+
- 🔄 Wait for ImageNet validation results
634

735
---
836

@@ -12,9 +40,11 @@ Systematic study of BitNet b1.58 (1.58-bit ternary quantization) applied to stan
1240

1341
**Core Narrative**: "For ternary CNNs, invest in distillation not augmentation."
1442

43+
**Practical Recipe**: Keep conv1 in FP32 + apply KD → recovers 88% of accuracy gap on CIFAR-10, **exceeds FP32 by 1.0% on CIFAR-100**.
44+
1545
---
1646

17-
## Current State (134 experiments completed)
47+
## Current State (153 experiments completed)
1848

1949
### Results Summary
2050

@@ -88,7 +118,38 @@ Systematic study of BitNet b1.58 (1.58-bit ternary quantization) applied to stan
88118

89119
**Status**: ✅ Complete (6 runs: 2 datasets × 3 seeds)
90120

91-
### Key Finding #5: ImageNet Validation 🔄 IN PROGRESS
121+
### Key Finding #5: conv1 + KD Combo ⭐ COMPLETE
122+
123+
**Combining conv1 in FP32 with KD recovers 88-123% of the accuracy gap.**
124+
125+
**CIFAR-10 Results (ResNet18)**:
126+
| Method | Accuracy | Gap Recovery |
127+
|--------|----------|--------------|
128+
| FP32 (target) | 88.89% | - |
129+
| BitNet (baseline) | 85.40% | 0% |
130+
| BitNet + KD | 86.66 ± 0.18% | 36% |
131+
| BitNet + keep_conv1 | 87.40% | 58% |
132+
| **BitNet + keep_conv1 + KD** | **88.48 ± 0.17%** | **88%** |
133+
134+
**CIFAR-100 Results (ResNet18)** ⭐ NEW:
135+
| Method | Accuracy | Gap Recovery |
136+
|--------|----------|--------------|
137+
| FP32 (target) | 62.40% | - |
138+
| BitNet (baseline) | 58.06% | 0% |
139+
| BitNet + KD | 60.55 ± 0.21% | 57% |
140+
| BitNet + keep_conv1 | 61.27% | 74% |
141+
| **BitNet + keep_conv1 + KD** | **63.40 ± 0.09%** | **123% (exceeds FP32!)** |
142+
143+
**Key insights**:
144+
145+
- Benefits are largely additive: conv1 (58%) + KD (36%) = 94% theoretical, achieved 88% on CIFAR-10
146+
- On CIFAR-100, the combo **exceeds FP32 by 1.0 percentage point** (63.40% vs 62.40%)
147+
- KD provides stronger regularization on harder tasks, pushing beyond FP32 baseline
148+
- This is the **practical recipe**: keep conv1 in FP32 + apply KD
149+
150+
**Status**: ✅ Complete (6 runs: ResNet18 × CIFAR-10/100 × seeds 42, 123, 456)
151+
152+
### Key Finding #6: ImageNet Validation 🔄 IN PROGRESS
92153

93154
**Goal**: Validate that findings scale beyond CIFAR to large-scale datasets.
94155

@@ -104,11 +165,11 @@ Systematic study of BitNet b1.58 (1.58-bit ternary quantization) applied to stan
104165

105166
**YES**: "Systematic analysis of BitNet scaling to standard deep CNN architectures (11M-25M params), revealing limitations not apparent in smaller networks"
106167

107-
### Paper Framing Options
168+
### Paper Framing ✅ FINALIZED
169+
170+
**Title**: "When Augmentation Fails: Knowledge Distillation for Ternary CNNs"
108171

109-
1. "The Augmentation Paradox: Why Data Augmentation Fails to Help Ternary Neural Networks"
110-
2. "BitNet b1.58 for CNNs: A Comprehensive Evaluation and Layer-Wise Analysis"
111-
3. "When 1.58 Bits Aren't Enough: Understanding the Limits of Extreme CNN Quantization"
172+
**Core message**: Standard data augmentation doesn't help ternary CNNs, but conv1 in FP32 + knowledge distillation recovers 88% of the accuracy gap on CIFAR-10 and **exceeds FP32 accuracy on CIFAR-100**.
112173

113174
---
114175

@@ -121,7 +182,7 @@ Systematic study of BitNet b1.58 (1.58-bit ternary quantization) applied to stan
121182
| **Layer-wise ablation** (quantize all-but-one layer) | 2-3 days | **High** | ✅ DONE |
122183
| **Efficiency proxy metrics** (BOPs, memory, model size) | Hours | Medium | ✅ DONE |
123184
| **KD experiment** (FP32 teacher → BitNet student) | 1-2 weeks | **High** | ✅ DONE (36% gap recovery) |
124-
| **conv1 + KD combo** (practical recipe) | ~6h | **High** | 🎯 PRIORITY |
185+
| **conv1 + KD combo** (practical recipe) | ~6h | **High** | ✅ DONE (88% gap recovery) |
125186
| **TTQ comparison** (address in Related Work) | Text only | **High** | 📝 Write explanation |
126187

127188
### Tier 2: Strengthens Paper
@@ -130,17 +191,35 @@ Systematic study of BitNet b1.58 (1.58-bit ternary quantization) applied to stan
130191
|------------|--------|--------|--------|
131192
| ImageNet validation (ResNet-18, 2 seeds) | 24-48h each | Medium | 🔄 Running |
132193
| KD on CIFAR-100 (harder task) | ~6h | Medium | ✅ DONE (57% gap recovery) |
194+
| **conv1 + KD on CIFAR-100** | ~6h | **High** | ✅ DONE (123% gap recovery - exceeds FP32!) |
133195

134196
### Tier 3: Optional (If Time Permits)
135197

136198
| Experiment | Effort | Signal | Status |
137199
|------------|--------|--------|--------|
138-
| KD on ResNet50/CIFAR-10 | ~12h | Low | Not started |
139-
| KD on ResNet50/CIFAR-100 | ~12h | Low | Not started |
200+
| KD on ResNet50/CIFAR-10 | ~12h | Low | ✅ DONE (77% recovery) |
201+
| KD on ResNet50/CIFAR-100 | ~12h | Low | ✅ DONE (81% recovery) |
140202
| Implement TTQ/DoReFa for direct comparison | 1-2 weeks | Medium | Not started |
141203
| Real inference measurements (BitBLAS) | Days | Medium | Not started |
142204
| Additional architectures (MobileNetV2) | Days | Low | Not started |
143205

206+
### Key Finding #7: ResNet50 Recipe Validation ⭐ NEW
207+
208+
**The conv1+KD recipe generalizes to larger models, though with lower recovery.**
209+
210+
| Model | Dataset | Result | FP32 | Gap Recovery |
211+
|-------|---------|--------|------|--------------|
212+
| ResNet50 | CIFAR-10 | 89.46% | 90.42% | **77%** |
213+
| ResNet50 | CIFAR-100 | 63.87% | 65.58% | **81%** |
214+
215+
**Comparison with ResNet18:**
216+
- ResNet18 CIFAR-10: 88% recovery
217+
- ResNet18 CIFAR-100: 123% recovery (exceeds FP32!)
218+
- ResNet50 CIFAR-10: 77% recovery
219+
- ResNet50 CIFAR-100: 81% recovery
220+
221+
**Insight**: Larger models are harder to fully recover. More parameters = more information lost to quantization. ResNet18 remains the sweet spot for ternary deployment.
222+
144223
### TTQ Comparison Strategy
145224

146225
**Don't implement TTQ.** Instead, explain the design tradeoff in Related Work:
@@ -153,32 +232,82 @@ This reframes the gap as an **intentional design tradeoff** for deployment effic
153232

154233
## Immediate Next Steps
155234

156-
### 🎯 PRIORITY 1: conv1 + KD Combo Experiment
235+
### 🎯 PRIORITY 1: Deep Literature Research ✅ COMPLETE
236+
237+
**Status**: All 4 research prompts completed and integrated into paper.
238+
239+
| File | Topic | Key Finding |
240+
|------|-------|-------------|
241+
| `01_ttq_comparison.md` | TTQ vs BitNet gap | TTQ uses learned asymmetric scales (W^p, W^n) + keeps conv1/FC in FP32. Our gap explained by simpler formulation. |
242+
| `02_layer_sensitivity_literature.md` | Layer sensitivity | conv1 FP32 is standard since 2016. **Our contribution: precise quantification (54-74%)**. |
243+
| `03_kd_for_quantization.md` | KD literature | T=4 may be suboptimal (T=6-8 better). **Feature distillation could add 5-20% more recovery**. |
244+
| `04_bitnet_cnn_prior_work.md` | Prior BitNet CNN work | Novelty confirmed: first ResNet study with full training + augmentation analysis. |
245+
246+
**Paper updates made**:
247+
- ✅ Related Work rewritten with proper framing and 6 new citations
248+
- ✅ Introduction updated to acknowledge conv1 FP32 as established practice
249+
- ✅ Contributions list refined ("quantified layer sensitivity" not "discovery")
250+
- ✅ Discussion added TTQ comparison paragraph (design tradeoff)
251+
- ✅ Future Work expanded with feature distillation and temperature tuning
252+
253+
### 🎯 PRIORITY 2: New Experiments from Research Findings
254+
255+
Based on KD research, these experiments could further improve results:
256+
257+
#### 2a. Higher Temperature KD ✅ COMPLETE
258+
259+
**Results**: T=4 is already optimal. Higher temperatures perform slightly worse.
260+
261+
| Temperature | Accuracy |
262+
|-------------|----------|
263+
| T=4 (default) | **88.66%** |
264+
| T=6 | 88.23% |
265+
| T=8 | 88.34% |
266+
267+
**Conclusion**: No benefit from higher temperatures. Our default T=4 is optimal for ternary networks.
268+
269+
#### 2b. CIFAR-100 conv1+KD (Validates Recipe) ✅ COMPLETE
270+
271+
**Results**: 63.40 ± 0.09% accuracy - **exceeds FP32 (62.40%) by 1.0 percentage point!**
272+
273+
| Seed | Accuracy |
274+
|------|----------|
275+
| 42 | 63.41% |
276+
| 123 | 63.48% |
277+
| 456 | 63.30% |
278+
279+
**Gap recovery**: 123% (exceeds 100% because KD regularization pushes beyond FP32)
280+
281+
**Key insight**: On harder tasks, KD provides stronger regularization, enabling ternary networks to surpass full-precision baselines.
157282

158-
**Goal**: Test if combining conv1 in FP32 with KD gives additive benefits.
283+
#### 2c. Feature Distillation (Bigger Effort, Bigger Gain)
284+
**Goal**: Implement DCQ/OFF-style feature distillation.
159285

160-
**Hypothesis**: conv1 (58% recovery) + KD (36% recovery) could recover ~85-90% of the gap.
286+
**Potential gain**: 5-20% additional improvement (research consensus).
161287

162-
**Method**:
288+
**Implementation**:
289+
- Add intermediate feature extraction to teacher/student
290+
- Cosine similarity loss on conv layer outputs
291+
- Combined loss: α×KD_logits + (1-α)×CE + β×feature_loss
163292

164-
- Add `--ablation keep_conv1` support to `train_kd.py`
165-
- Run 3 seeds on ResNet18/CIFAR-10
166-
- Compare: baseline → +conv1 → +KD → +both
293+
**Status**: Parked for now. Could push from 88% to ~95% recovery. Good for future work or camera-ready revision.
167294

168-
**Expected outcome**: Practical recipe for maximum accuracy with minimal overhead.
295+
### 🎯 PRIORITY 3: Paper Polish
169296

170-
### 🎯 PRIORITY 2: TTQ Explanation in Paper
297+
**Goal**: Final polish before submission.
171298

172-
**Goal**: Address why BitNet underperforms TTQ (which claimed to beat FP32).
299+
**Tasks**:
173300

174-
**Approach**: Explain in Related Work as design tradeoff, not implement.
301+
- [ ] Generate Fig 4 (Recipe comparison bar chart)
302+
- [ ] Verify paper compiles with new citations
303+
- [ ] Final proofread
304+
- [ ] Wait for ImageNet validation results
175305

176306
---
177307

178308
### Currently Running
179309

180310
- 🔄 **ImageNet validation**: 4 runs (ResNet18, seeds 42/123, FP32/BitNet)
181-
- ✅ **KD on CIFAR-100**: 60.55 ± 0.21% (+2.49%, 57% gap recovery) - **KD more effective on harder tasks!**
182311

183312
---
184313

@@ -187,8 +316,15 @@ This reframes the gap as an **intentional design tradeoff** for deployment effic
187316
- ✅ **Layer-wise ablation**: conv1 accounts for 54-74% of accuracy gap (confirmed across all configs)
188317
- ✅ **Ablation generalization**: Tested on ResNet18/50 × CIFAR-10/100 (36 ablation runs)
189318
- ✅ **Efficiency metrics**: 20.3x compression, 64x theoretical speedup
190-
- ✅ **134 total experiments**: Main experiments + comprehensive ablation study
191-
- ✅ **Knowledge Distillation**: CIFAR-10 (36% recovery), CIFAR-100 (57% recovery) - KD more effective on harder tasks
319+
- ✅ **Knowledge Distillation**: CIFAR-10 (36% recovery), CIFAR-100 (57% recovery)
320+
- ✅ **conv1 + KD combo (CIFAR-10)**: 88.48 ± 0.17% accuracy (88% gap recovery)
321+
- ✅ **conv1 + KD combo (CIFAR-100)**: 63.40 ± 0.09% accuracy (**exceeds FP32 by 1.0%!**)
322+
- ✅ **Temperature ablation**: T=4 is optimal, T=6/T=8 perform worse
323+
- ✅ **ResNet50 recipe validation**: 77-81% gap recovery (lower than ResNet18 but still substantial)
324+
- ✅ **153 total experiments**: Main + ablation + KD studies
325+
- ✅ **Paper first draft**: All main sections written (13 pages)
326+
- ✅ **Research prompts**: Created 4 deep research prompts for AI collaboration
327+
- ✅ **Reproducibility appendix**: Documented code → results → paper pipeline
192328

193329
---
194330

@@ -209,12 +345,13 @@ This reframes the gap as an **intentional design tradeoff** for deployment effic
209345

210346
| Month | Tasks |
211347
|-------|-------|
212-
| **Jan-Feb** | ✅ Layer-wise ablation, efficiency metrics, KD experiment |
213-
| **Feb** | conv1+KD combo, ImageNet results, CIFAR-100 KD |
214-
| **Mar-Apr** | Paper writing, figures, internal review |
215-
| **May-Jun** | Revisions, polish |
216-
| **Jul-Aug** | Submit to NeurIPS 2026 Efficient ML Workshop |
217-
| **Sept** | Submit to WACV 2027 Round 2 |
348+
| **Jan** | ✅ Layer-wise ablation, efficiency metrics, KD experiment |
349+
| **Feb (now)** | ✅ conv1+KD combo, ✅ Paper first draft, 🔄 ImageNet, 📝 CIFAR-100 conv1+KD |
350+
| **Feb-Mar** | Deep research integration, figures, CIFAR-100 conv1+KD results |
351+
| **Mar** | Polish, internal review |
352+
| **~Apr** | Submit to **CVPR 2026 Workshop** (primary target) |
353+
| **Aug** | Submit to NeurIPS 2026 Efficient ML Workshop (if needed) |
354+
| **Sept** | Submit to WACV 2027 Round 2 (backup) |
218355

219356
---
220357

@@ -229,8 +366,9 @@ This reframes the gap as an **intentional design tradeoff** for deployment effic
229366

230367
| File | Purpose |
231368
|------|---------|
369+
| `paper/main.tex` | LaTeX paper (first draft complete) |
232370
| `paper/notes.md` | Detailed paper notes and TODO |
233-
| `paper/main.tex` | LaTeX paper skeleton |
371+
| `paper/research/` | Deep research prompts for AI collaboration |
234372
| `paper/prompts_for_review.md` | Prompts for external review |
235373
| `results/processed/aggregated.csv` | Current experiment results |
236374

0 commit comments

Comments
 (0)