Repository navigation
test: cover image transform helpers used by Idefics3 and Jina CLIP preprocessing - #785
Aditya17-bot wants to merge 1 commit into
Conversation
…eprocessing pad2square, crop_ndarray, resize_ndarray, ResizeForVisionEncoder, ImageSplitter and SquareResize had no unit tests; they were only reached through model tests that download weights. Expected sizes follow the Idefics3 image processor in transformers.
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (1)
Included review availability: This review used your included allowance. Your plan provides up to 8 included reviews per hour; 7 remain after this review. 📝 WalkthroughWalkthroughThe change adds tests for image padding, cropping, and resizing, including channel layouts and output dtype expectations. It adds assertions for vision-encoder dimensions, image-splitting tile contents and counts, per-image grouping, and Priority: ⬇️ Low Estimated code review effort: 2 (Simple) | ~12 minutes Change: Other Merge Risk: ⚪ Minimal · up to No merge-blocking issue is identified; merge after normal checks. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Adds unit tests for the image transform helpers that only ran through model tests (which download weights):
pad2square,crop_ndarray,resize_ndarray,ResizeForVisionEncoder,ImageSplitterandSquareResize. They are used by Jina CLIP (PadtoSquare) and ColModernVBERT (Idefics3 resizing and splitting).What the tests pin:
pad2square: a smaller image is pasted top-left on the fill color; a larger one is center-cropped with round-half-to-even offsets (7 -> 2, 9 -> 2), matching the torchvision note in the code.crop_ndarray: channel-first and channel-last crops agree.resize_ndarray:sizeis(width, height), the layout stays(C, H, W), uint8 stays uint8, and float images stay in [0, 1].ResizeForVisionEncoder: landscape, portrait, smaller-than-one-tile and exact-tile inputs round up to multiples of the tile size, with the same sizes asresize_for_vision_encoderin transformers' Idefics3 image processor.ImageSplitter: small images pass through untouched; larger ones giverows * colstiles in row order plus a 512x512 global view; per-image grouping is kept.SquareResize: each image is wrapped in its own list atsize x size.ImageSplitteris only tested on multiples of the tile size, sinceResizeForVisionEncoderalways runs first. For other sizes it splits into evenly sized patches, whereas transformers' currentsplit_imagestakes fixed tiles, so I didn't want to pin that path.The tests are in a new file so they don't conflict with #760 and #761, which both edit
tests/test_image_transform.py. I leftresize_longest_edgeout because #761 covers it.Coverage from
tests/test_image_transform.pyalone vs. with the new file:functional.py: 57% -> 82%operators.py: 43% -> 62%