11# BitNet CNN Research Plan
22
3- ** Last Updated** : Feb 2 , 2026
4- ** Primary Target** : WACV 2027 Round 2 (Sept 2026)
3+ ** Last Updated** : Feb 5 , 2026
4+ ** Primary Target** : CVPR 2026 Workshop ( ~ Apr 2026)
55** Parallel Track** : NeurIPS 2026 Efficient ML Workshop (Aug 2026)
6+ ** Backup** : WACV 2027 Round 2 (Sept 2026)
7+
8+ ---
9+
10+ ## Paper Status: 📝 FIRST DRAFT COMPLETE
11+
12+ ** Title** : "When Augmentation Fails: Knowledge Distillation for Ternary CNNs"
13+
14+ ** Paper file** : ` paper/main.tex ` (13 pages, builds successfully)
15+
16+ ### Sections Completed
17+
18+ - ✅ Title and author info (Dario Cazzani, Cisco Systems Inc.)
19+ - ✅ Abstract
20+ - ✅ Introduction (framing the augmentation paradox)
21+ - ✅ Related Work (with TTQ, QKD, HAQ citations)
22+ - ✅ Method (BitNet b1.58 implementation details)
23+ - ✅ Section 4: "What Doesn't Work: The Augmentation Paradox"
24+ - ✅ Section 5: "What Works: A Practical Recipe" (conv1 + KD)
25+ - ✅ Discussion
26+ - ✅ Conclusion
27+ - ✅ Reproducibility Appendix (code → results → paper pipeline)
28+
29+ ### Still Needed
30+
31+ - 📝 Complete figures (Fig 4: Recipe comparison bar chart)
32+ - 📝 Polish and proofread
33+ - 🔄 Wait for ImageNet validation results
634
735---
836
@@ -12,9 +40,11 @@ Systematic study of BitNet b1.58 (1.58-bit ternary quantization) applied to stan
1240
1341** Core Narrative** : "For ternary CNNs, invest in distillation not augmentation."
1442
43+ ** Practical Recipe** : Keep conv1 in FP32 + apply KD → recovers 88% of accuracy gap on CIFAR-10, ** exceeds FP32 by 1.0% on CIFAR-100** .
44+
1545---
1646
17- ## Current State (134 experiments completed)
47+ ## Current State (153 experiments completed)
1848
1949### Results Summary
2050
@@ -88,7 +118,38 @@ Systematic study of BitNet b1.58 (1.58-bit ternary quantization) applied to stan
88118
89119** Status** : ✅ Complete (6 runs: 2 datasets × 3 seeds)
90120
91- ### Key Finding #5 : ImageNet Validation 🔄 IN PROGRESS
121+ ### Key Finding #5 : conv1 + KD Combo ⭐ COMPLETE
122+
123+ ** Combining conv1 in FP32 with KD recovers 88-123% of the accuracy gap.**
124+
125+ ** CIFAR-10 Results (ResNet18)** :
126+ | Method | Accuracy | Gap Recovery |
127+ | --------| ----------| --------------|
128+ | FP32 (target) | 88.89% | - |
129+ | BitNet (baseline) | 85.40% | 0% |
130+ | BitNet + KD | 86.66 ± 0.18% | 36% |
131+ | BitNet + keep_conv1 | 87.40% | 58% |
132+ | ** BitNet + keep_conv1 + KD** | ** 88.48 ± 0.17%** | ** 88%** |
133+
134+ ** CIFAR-100 Results (ResNet18)** ⭐ NEW:
135+ | Method | Accuracy | Gap Recovery |
136+ | --------| ----------| --------------|
137+ | FP32 (target) | 62.40% | - |
138+ | BitNet (baseline) | 58.06% | 0% |
139+ | BitNet + KD | 60.55 ± 0.21% | 57% |
140+ | BitNet + keep_conv1 | 61.27% | 74% |
141+ | ** BitNet + keep_conv1 + KD** | ** 63.40 ± 0.09%** | ** 123% (exceeds FP32!)** |
142+
143+ ** Key insights** :
144+
145+ - Benefits are largely additive: conv1 (58%) + KD (36%) = 94% theoretical, achieved 88% on CIFAR-10
146+ - On CIFAR-100, the combo ** exceeds FP32 by 1.0 percentage point** (63.40% vs 62.40%)
147+ - KD provides stronger regularization on harder tasks, pushing beyond FP32 baseline
148+ - This is the ** practical recipe** : keep conv1 in FP32 + apply KD
149+
150+ ** Status** : ✅ Complete (6 runs: ResNet18 × CIFAR-10/100 × seeds 42, 123, 456)
151+
152+ ### Key Finding #6 : ImageNet Validation 🔄 IN PROGRESS
92153
93154** Goal** : Validate that findings scale beyond CIFAR to large-scale datasets.
94155
@@ -104,11 +165,11 @@ Systematic study of BitNet b1.58 (1.58-bit ternary quantization) applied to stan
104165
105166** YES** : "Systematic analysis of BitNet scaling to standard deep CNN architectures (11M-25M params), revealing limitations not apparent in smaller networks"
106167
107- ### Paper Framing Options
168+ ### Paper Framing ✅ FINALIZED
169+
170+ ** Title** : "When Augmentation Fails: Knowledge Distillation for Ternary CNNs"
108171
109- 1 . "The Augmentation Paradox: Why Data Augmentation Fails to Help Ternary Neural Networks"
110- 2 . "BitNet b1.58 for CNNs: A Comprehensive Evaluation and Layer-Wise Analysis"
111- 3 . "When 1.58 Bits Aren't Enough: Understanding the Limits of Extreme CNN Quantization"
172+ ** Core message** : Standard data augmentation doesn't help ternary CNNs, but conv1 in FP32 + knowledge distillation recovers 88% of the accuracy gap on CIFAR-10 and ** exceeds FP32 accuracy on CIFAR-100** .
112173
113174---
114175
@@ -121,7 +182,7 @@ Systematic study of BitNet b1.58 (1.58-bit ternary quantization) applied to stan
121182| ** Layer-wise ablation** (quantize all-but-one layer) | 2-3 days | ** High** | ✅ DONE |
122183| ** Efficiency proxy metrics** (BOPs, memory, model size) | Hours | Medium | ✅ DONE |
123184| ** KD experiment** (FP32 teacher → BitNet student) | 1-2 weeks | ** High** | ✅ DONE (36% gap recovery) |
124- | ** conv1 + KD combo** (practical recipe) | ~ 6h | ** High** | 🎯 PRIORITY |
185+ | ** conv1 + KD combo** (practical recipe) | ~ 6h | ** High** | ✅ DONE (88% gap recovery) |
125186| ** TTQ comparison** (address in Related Work) | Text only | ** High** | 📝 Write explanation |
126187
127188### Tier 2: Strengthens Paper
@@ -130,17 +191,35 @@ Systematic study of BitNet b1.58 (1.58-bit ternary quantization) applied to stan
130191| ------------| --------| --------| --------|
131192| ImageNet validation (ResNet-18, 2 seeds) | 24-48h each | Medium | 🔄 Running |
132193| KD on CIFAR-100 (harder task) | ~ 6h | Medium | ✅ DONE (57% gap recovery) |
194+ | ** conv1 + KD on CIFAR-100** | ~ 6h | ** High** | ✅ DONE (123% gap recovery - exceeds FP32!) |
133195
134196### Tier 3: Optional (If Time Permits)
135197
136198| Experiment | Effort | Signal | Status |
137199| ------------| --------| --------| --------|
138- | KD on ResNet50/CIFAR-10 | ~ 12h | Low | Not started |
139- | KD on ResNet50/CIFAR-100 | ~ 12h | Low | Not started |
200+ | KD on ResNet50/CIFAR-10 | ~ 12h | Low | ✅ DONE (77% recovery) |
201+ | KD on ResNet50/CIFAR-100 | ~ 12h | Low | ✅ DONE (81% recovery) |
140202| Implement TTQ/DoReFa for direct comparison | 1-2 weeks | Medium | Not started |
141203| Real inference measurements (BitBLAS) | Days | Medium | Not started |
142204| Additional architectures (MobileNetV2) | Days | Low | Not started |
143205
206+ ### Key Finding #7 : ResNet50 Recipe Validation ⭐ NEW
207+
208+ ** The conv1+KD recipe generalizes to larger models, though with lower recovery.**
209+
210+ | Model | Dataset | Result | FP32 | Gap Recovery |
211+ | -------| ---------| --------| ------| --------------|
212+ | ResNet50 | CIFAR-10 | 89.46% | 90.42% | ** 77%** |
213+ | ResNet50 | CIFAR-100 | 63.87% | 65.58% | ** 81%** |
214+
215+ ** Comparison with ResNet18:**
216+ - ResNet18 CIFAR-10: 88% recovery
217+ - ResNet18 CIFAR-100: 123% recovery (exceeds FP32!)
218+ - ResNet50 CIFAR-10: 77% recovery
219+ - ResNet50 CIFAR-100: 81% recovery
220+
221+ ** Insight** : Larger models are harder to fully recover. More parameters = more information lost to quantization. ResNet18 remains the sweet spot for ternary deployment.
222+
144223### TTQ Comparison Strategy
145224
146225** Don't implement TTQ.** Instead, explain the design tradeoff in Related Work:
@@ -153,32 +232,82 @@ This reframes the gap as an **intentional design tradeoff** for deployment effic
153232
154233## Immediate Next Steps
155234
156- ### 🎯 PRIORITY 1: conv1 + KD Combo Experiment
235+ ### 🎯 PRIORITY 1: Deep Literature Research ✅ COMPLETE
236+
237+ ** Status** : All 4 research prompts completed and integrated into paper.
238+
239+ | File | Topic | Key Finding |
240+ | ------| -------| -------------|
241+ | ` 01_ttq_comparison.md ` | TTQ vs BitNet gap | TTQ uses learned asymmetric scales (W^p, W^n) + keeps conv1/FC in FP32. Our gap explained by simpler formulation. |
242+ | ` 02_layer_sensitivity_literature.md ` | Layer sensitivity | conv1 FP32 is standard since 2016. ** Our contribution: precise quantification (54-74%)** . |
243+ | ` 03_kd_for_quantization.md ` | KD literature | T=4 may be suboptimal (T=6-8 better). ** Feature distillation could add 5-20% more recovery** . |
244+ | ` 04_bitnet_cnn_prior_work.md ` | Prior BitNet CNN work | Novelty confirmed: first ResNet study with full training + augmentation analysis. |
245+
246+ ** Paper updates made** :
247+ - ✅ Related Work rewritten with proper framing and 6 new citations
248+ - ✅ Introduction updated to acknowledge conv1 FP32 as established practice
249+ - ✅ Contributions list refined ("quantified layer sensitivity" not "discovery")
250+ - ✅ Discussion added TTQ comparison paragraph (design tradeoff)
251+ - ✅ Future Work expanded with feature distillation and temperature tuning
252+
253+ ### 🎯 PRIORITY 2: New Experiments from Research Findings
254+
255+ Based on KD research, these experiments could further improve results:
256+
257+ #### 2a. Higher Temperature KD ✅ COMPLETE
258+
259+ ** Results** : T=4 is already optimal. Higher temperatures perform slightly worse.
260+
261+ | Temperature | Accuracy |
262+ | -------------| ----------|
263+ | T=4 (default) | ** 88.66%** |
264+ | T=6 | 88.23% |
265+ | T=8 | 88.34% |
266+
267+ ** Conclusion** : No benefit from higher temperatures. Our default T=4 is optimal for ternary networks.
268+
269+ #### 2b. CIFAR-100 conv1+KD (Validates Recipe) ✅ COMPLETE
270+
271+ ** Results** : 63.40 ± 0.09% accuracy - ** exceeds FP32 (62.40%) by 1.0 percentage point!**
272+
273+ | Seed | Accuracy |
274+ | ------| ----------|
275+ | 42 | 63.41% |
276+ | 123 | 63.48% |
277+ | 456 | 63.30% |
278+
279+ ** Gap recovery** : 123% (exceeds 100% because KD regularization pushes beyond FP32)
280+
281+ ** Key insight** : On harder tasks, KD provides stronger regularization, enabling ternary networks to surpass full-precision baselines.
157282
158- ** Goal** : Test if combining conv1 in FP32 with KD gives additive benefits.
283+ #### 2c. Feature Distillation (Bigger Effort, Bigger Gain)
284+ ** Goal** : Implement DCQ/OFF-style feature distillation.
159285
160- ** Hypothesis ** : conv1 (58% recovery) + KD (36% recovery) could recover ~ 85-90% of the gap .
286+ ** Potential gain ** : 5-20% additional improvement (research consensus) .
161287
162- ** Method** :
288+ ** Implementation** :
289+ - Add intermediate feature extraction to teacher/student
290+ - Cosine similarity loss on conv layer outputs
291+ - Combined loss: α×KD_logits + (1-α)×CE + β×feature_loss
163292
164- - Add ` --ablation keep_conv1 ` support to ` train_kd.py `
165- - Run 3 seeds on ResNet18/CIFAR-10
166- - Compare: baseline → +conv1 → +KD → +both
293+ ** Status** : Parked for now. Could push from 88% to ~ 95% recovery. Good for future work or camera-ready revision.
167294
168- ** Expected outcome ** : Practical recipe for maximum accuracy with minimal overhead.
295+ ### 🎯 PRIORITY 3: Paper Polish
169296
170- ### 🎯 PRIORITY 2: TTQ Explanation in Paper
297+ ** Goal ** : Final polish before submission.
171298
172- ** Goal ** : Address why BitNet underperforms TTQ (which claimed to beat FP32).
299+ ** Tasks ** :
173300
174- ** Approach** : Explain in Related Work as design tradeoff, not implement.
301+ - [ ] Generate Fig 4 (Recipe comparison bar chart)
302+ - [ ] Verify paper compiles with new citations
303+ - [ ] Final proofread
304+ - [ ] Wait for ImageNet validation results
175305
176306---
177307
178308### Currently Running
179309
180310- 🔄 ** ImageNet validation** : 4 runs (ResNet18, seeds 42/123, FP32/BitNet)
181- - ✅ ** KD on CIFAR-100** : 60.55 ± 0.21% (+2.49%, 57% gap recovery) - ** KD more effective on harder tasks!**
182311
183312---
184313
@@ -187,8 +316,15 @@ This reframes the gap as an **intentional design tradeoff** for deployment effic
187316- ✅ ** Layer-wise ablation** : conv1 accounts for 54-74% of accuracy gap (confirmed across all configs)
188317- ✅ ** Ablation generalization** : Tested on ResNet18/50 × CIFAR-10/100 (36 ablation runs)
189318- ✅ ** Efficiency metrics** : 20.3x compression, 64x theoretical speedup
190- - ✅ ** 134 total experiments** : Main experiments + comprehensive ablation study
191- - ✅ ** Knowledge Distillation** : CIFAR-10 (36% recovery), CIFAR-100 (57% recovery) - KD more effective on harder tasks
319+ - ✅ ** Knowledge Distillation** : CIFAR-10 (36% recovery), CIFAR-100 (57% recovery)
320+ - ✅ ** conv1 + KD combo (CIFAR-10)** : 88.48 ± 0.17% accuracy (88% gap recovery)
321+ - ✅ ** conv1 + KD combo (CIFAR-100)** : 63.40 ± 0.09% accuracy (** exceeds FP32 by 1.0%!** )
322+ - ✅ ** Temperature ablation** : T=4 is optimal, T=6/T=8 perform worse
323+ - ✅ ** ResNet50 recipe validation** : 77-81% gap recovery (lower than ResNet18 but still substantial)
324+ - ✅ ** 153 total experiments** : Main + ablation + KD studies
325+ - ✅ ** Paper first draft** : All main sections written (13 pages)
326+ - ✅ ** Research prompts** : Created 4 deep research prompts for AI collaboration
327+ - ✅ ** Reproducibility appendix** : Documented code → results → paper pipeline
192328
193329---
194330
@@ -209,12 +345,13 @@ This reframes the gap as an **intentional design tradeoff** for deployment effic
209345
210346| Month | Tasks |
211347| -------| -------|
212- | ** Jan-Feb** | ✅ Layer-wise ablation, efficiency metrics, KD experiment |
213- | ** Feb** | conv1+KD combo, ImageNet results, CIFAR-100 KD |
214- | ** Mar-Apr** | Paper writing, figures, internal review |
215- | ** May-Jun** | Revisions, polish |
216- | ** Jul-Aug** | Submit to NeurIPS 2026 Efficient ML Workshop |
217- | ** Sept** | Submit to WACV 2027 Round 2 |
348+ | ** Jan** | ✅ Layer-wise ablation, efficiency metrics, KD experiment |
349+ | ** Feb (now)** | ✅ conv1+KD combo, ✅ Paper first draft, 🔄 ImageNet, 📝 CIFAR-100 conv1+KD |
350+ | ** Feb-Mar** | Deep research integration, figures, CIFAR-100 conv1+KD results |
351+ | ** Mar** | Polish, internal review |
352+ | ** ~ Apr** | Submit to ** CVPR 2026 Workshop** (primary target) |
353+ | ** Aug** | Submit to NeurIPS 2026 Efficient ML Workshop (if needed) |
354+ | ** Sept** | Submit to WACV 2027 Round 2 (backup) |
218355
219356---
220357
@@ -229,8 +366,9 @@ This reframes the gap as an **intentional design tradeoff** for deployment effic
229366
230367| File | Purpose |
231368| ------| ---------|
369+ | ` paper/main.tex ` | LaTeX paper (first draft complete) |
232370| ` paper/notes.md ` | Detailed paper notes and TODO |
233- | ` paper/main.tex ` | LaTeX paper skeleton |
371+ | ` paper/research/ ` | Deep research prompts for AI collaboration |
234372| ` paper/prompts_for_review.md ` | Prompts for external review |
235373| ` results/processed/aggregated.csv ` | Current experiment results |
236374
0 commit comments