English | 简体中文
Complete reference for the bcdl Python module (the nanobind bindings over the
C++ core). For the C++ API see CPP_API.md (the public headers
under include/bcdl/ are the source of truth); the names map
1:1 (Python snake_case ⇄ C++ camelCase).
- Install · Two ways to use BCDL · Conventions
- Engine: Engine
- Preprocessing: Letterbox & geometry · NV12 helpers
- Tasks: Detection · Classification · Pose · Instance segmentation · Oriented boxes (OBB) · Semantic segmentation · Depth · Stereo · OCR
- Streaming & tracking: Tracking (ByteTrack) · TrackingPipeline · AsyncDetectionPipeline
- Media: Images & codecs
Naming: snake_case functions and
lower_caseconfig fields are the Python spelling of the C++ camelCase members — without exception. Nine decoders used to carry a camelCase alias as well; those names are deprecated, resolve with aDeprecationWarning, and are no longer part of this reference.
See the README for the full table. The fast path:
conda install -c https://mirrors.ruis.ai/conda -c conda-forge bcdl
python -c "import bcdl; print(bcdl.__version__)"import bcdl pulls the C++ library and the packaged hobot SDK transitively, so
nothing else is needed on an RDK S100 / S100P / S600 board.
Every task is available at two levels — pick per how much you want BCDL to own:
-
High-level task classes (
Detector,Classifier,PoseEstimator,Segmenter,DepthEstimator, …). You give them anEngineand a config; they run inference and post-process in one call. BCDL owns the device buffers, cache flushes, and dequantization. -
Pure
decode_*functions (decode,decode_pose,decode_obb,decode_seg,decode_ctc, …). These take float32 NumPy arrays (a model output you already have) and return the same result objects — no Engine, no board, no model needed. This is the path the deterministic tests use, and the way to plug BCDL's post-processing onto an output you produced elsewhere.Quantization caveat:
decode_*assume already-dequantized float32. For anF16output, cast first (np.asarray(out, np.float32)). For a quantized int8/int16 output they would read raw integers and produce wrong results — use the high-level class instead, which dequantizes via the tensor's scale/zero-point.
- Images are
HxWx3uint8 BGR (OpenCV order). Pipelines reject non-uint8 input loudly rather than silently detecting nothing. - Coordinates come back in original-image pixels. Tasks that need to undo
the letterbox take a
LetterboxInfoinpostprocess(...). - NV12 two-input models: the standard RDK YOLO export takes two inputs
[Y, UV]; the*Detector.detect([...], lb)wrappers set inputs in order. - Preprocessing ownership: the high-level task classes (
Detector, etc.) leave preprocessing to you (model layouts vary). The pipelines (TrackingPipeline,AsyncDetectionPipeline,StereoPipeline) own preprocessing in C++ and take a raw BGR frame. __version__isbcdl.__version__.
Loads a compiled .hbm and runs BPU inference, handling the cache discipline
(clean inputs before infer, invalidate outputs before read) for you.
import numpy as np, bcdl
engine = bcdl.Engine("model.hbm") # optional model_name=""
print(engine.num_inputs, engine.num_outputs)
print(engine.input_shape(0), engine.input_dtype(0))
print(engine.output_shape(0), engine.output_dtype(0))
# Generic infer: list of input arrays in -> list of output arrays out.
outs = engine.infer([x]) # outs[i] reshaped to output_shape(i)| member | description |
|---|---|
Engine(hbm_path, model_name="") |
Load a model. model_name selects one packaged model if the .hbm holds several. |
model_name -> str |
The active model's name. |
num_inputs / num_outputs -> int |
Tensor counts. |
input_shape(i) / output_shape(i) -> list[int] |
Tensor shape. |
input_bytes(i) -> int |
Allocated device-buffer size for input i (after BPU stride alignment). |
input_packed_bytes(i) -> int |
Size of input i as a contiguous row-major array. Smaller than input_bytes(i) when the model pads a dimension. |
input_stride(i) -> list[int] |
Resolved byte strides of input i, innermost last. |
output_stride(i) -> list[int] |
Resolved byte strides of output i. Outputs pad too — reshaping the raw buffer flat shears the tensor. |
output_packed_bytes(i) -> int |
Size of output i as a contiguous row-major array. Smaller than the device buffer when the model pads a dimension. |
input_dtype(i) / output_dtype(i) -> str |
NumPy dtype string (e.g. "float32", "float16", "int8"). |
infer(inputs, timeout_ms=0) -> list[np.ndarray] |
Copy inputs into device buffers, run, return one array per output. timeout_ms=0 blocks. |
Padded outputs are gathered for you.
infer()returns packed row-major arrays even when the model pads a dimension (ViTPose's[1,133,64,48]heatmap has a 256-byte row stride against 192 valid bytes). Reading the device buffer flat instead displaces rowrbyr * padand yields a tensor that still looks like data — it cost a day of debugging a model that was in fact correct.
The high-level task classes wrap an
Engineand call into the native post-processor directly (no per-call NumPy round-trip), so prefer them for a single task; useEngine.infer()when you want the raw output tensors.
Aspect-preserving fit and coordinate mapping between original and model space.
lb = bcdl.compute_letterbox(src_w, src_h, dst_w, dst_h, center_pad=True)
mx, my = lb.fwd_x(x), lb.fwd_y(y) # original -> model pixel
ox, oy = lb.inv_x(x), lb.inv_y(y) # model -> original pixel
# One-shot CPU letterbox of a uint8 image (needs OpenCV):
canvas, lb = bcdl.letterbox_numpy(img, dst_w, dst_h, pad=114)compute_letterbox(src_w, src_h, dst_w, dst_h, center_pad=True) -> LetterboxInfoLetterboxInfo— fieldsscale,pad_x,pad_y,src_w/h,dst_w/h(all read/write); methodsfwd_x/fwd_y(orig→model) andinv_x/inv_y(model→orig). Pass it to every detection-familypostprocess(...)so boxes come back in original pixels.letterbox_numpy(img, dst_w, dst_h, pad=114) -> (canvas, LetterboxInfo)— geometry matches the C++ VP letterbox; requires OpenCV.
nv12 = bcdl.bgr_to_nv12(bgr) # (H*3//2, W) uint8; even dims; needs OpenCVbgr_to_nv12(bgr) -> np.ndarray— packed NV12, feedsvp_image_from_nv12and the JPU encoder. Also see Images & codecs.
Fixed geometric transforms on the VPS GDC engine — NV12 in/out, CPU idle
during the op. None off-board / without GDC support. Full semantics + the
reverse-engineered CUSTOM-grid notes: docs/GDC.md.
# hardware letterbox (warp LUT generated at construction)
g = bcdl.GdcLetterbox(in_w, in_h, out_w, out_h, pad=114)
dst = g.run(src_vpimage) # NV12 VpImage -> NV12 VpImage
# hardware dense remap (cv2.remap semantics; LUT generated at runtime)
g = bcdl.GdcRemap(map_x, map_y, in_w, in_h, grid_step=16) # (out_h,out_w) f32 maps
dst = g.run(src_vpimage)GdcRemap(map_x, map_y, in_w, in_h, grid_step=16)— arbitrary FIXED warp,out(x,y) = in(map_x[y,x], map_y[y,x]); built for stereo rectification.grid_stepmust divide both output dims. 2448×2048: 6.3 ms wall / ~1 ms CPU vs 14.7 ms all-corescv2.remap; matches cv2 to p99 ≤ 2 grey-levels.BCDL_GDC_TIMING=1prints copy/op breakdown.GdcLetterbox(in_w, in_h, out_w, out_h, pad=114)— letterbox on GDC; the warp LUT is generated at construction (no offline.bin)..inforeturns theLetterboxInfo. 1920×1080 → 640×640: 0.97 ms wall, ~0.3 ms of it CPU. Geometry matches the CPU letterbox exactly; the resamplers alias differently on high-frequency detail — see docs/GDC.md.
Two decoder families:
- Single-tensor (
DecodeLayout.YoloV8/YoloV5) — one float output tensor. UseDetectConfig+decode(...)/Detector. - Anchor-free LTRB multi-scale (YOLO26 / standard RDK NV12 export) — a
(cls, box)output pair per stride. UseYoloLtrbConfig+YoloLtrbDetector.
# Single-tensor, numpy path:
cfg = bcdl.DetectConfig(); cfg.num_classes = 80; cfg.input_w = cfg.input_h = 640
out = engine.infer([x])[0].astype(np.float32) # (N, 4+nc) or (4+nc, N)
dets = bcdl.decode(out, cfg, lb)
for d in dets:
print(d.class_id, d.score, d.x1, d.y1, d.x2, d.y2)
# Anchor-free LTRB, high-level (NV12 two-input):
det = bcdl.YoloLtrbDetector(engine, bcdl.YoloLtrbConfig())
dets = det.detect([y_plane, uv_plane], lb)DetectConfig — input_w, input_h, num_classes, conf_thresh,
iou_thresh, max_dets, layout (DecodeLayout), channels_first,
apply_sigmoid.
DecodeLayout enum — YoloV8, YoloV5.
Detection — x1, y1, x2, y2, score, class_id (all read/write); __repr__.
YoloLtrbConfig — num_classes, conf_thresh, iou_thresh, max_dets,
strides (e.g. [8,16,32]), reg_max (DFL bins; 0/1 = direct LTRB).
Functions / classes:
| symbol | signature |
|---|---|
decode |
decode(output: f32 ndarray, config: DetectConfig, letterbox: LetterboxInfo) -> list[Detection] |
nms |
nms(dets, iou_thresh, max_dets=300) -> list[int] (indices to keep) |
iou |
iou(a: Detection, b: Detection) -> float |
Detector |
Detector(engine, config, output_index=0); .detect(model_input, lb, timeout_ms=0), .config |
YoloLtrbDetector |
YoloLtrbDetector(engine, config=None, output_base=0); .detect([inputs], lb, timeout_ms=0), .postprocess(lb), .config |
cfg = bcdl.ClsConfig(); cfg.top_k = 5; cfg.apply_softmax = True
clf = bcdl.Classifier(engine, cfg)
for r in clf.classify(x): # x: single array or [Y, UV]
print(r.class_id, r.score)
# numpy path:
top = bcdl.decode_classification(logits.astype(np.float32), cfg)ClsConfig—top_k,apply_softmax.ClsResult—class_id,score;__repr__.decode_classification(logits, config) -> list[ClsResult]— flat logit vector top-k.Classifier(engine, config=None, output_index=0)—.classify(inputs, timeout_ms=0),.postprocess(),.config.
LTRB multi-scale head: a person box plus K keypoints. NV12 two-input.
cfg = bcdl.PoseConfig(); cfg.num_keypoints = 17
est = bcdl.PoseEstimator(engine, cfg)
for p in est.detect([y_plane, uv_plane], lb):
print(p.score, [(k.x, k.y, k.score) for k in p.keypoints])PoseConfig—num_keypoints,conf_thresh,iou_thresh,max_dets,strides.PoseDetection—x1, y1, x2, y2, score, class_id,keypoints: list[Keypoint].Keypoint—x, y, score.PoseEstimator(engine, config=None, output_base=0)—.detect([inputs], lb, timeout_ms=0),.postprocess(lb),.config.decode_pose(cls, box, kpt, config, letterbox) -> list[PoseDetection]— numpy path;cls/box/kptare lists of per-stride[H,W,1]/[H,W,4]/[H,W,K*3]float arrays.
TOP-DOWN, unlike PoseEstimator: one inference per PERSON on a crop, so it needs
a detector in front of it and its cost scales with the head count. In exchange it
resolves feet, a 68-point face and both hands.
est = bcdl.WholeBodyEstimator(engine)
for box in person_boxes: # from any detector above
kpts = est.estimate(bgr, box) # 133 keypoints, original-image px
print(sum(k.score > 0.2 for k in kpts), "visible")Keypoint layout (COCO-WholeBody): 0-16 body, 17-22 feet, 23-90 face,
91-111 left hand, 112-132 right hand.
WholeBodyConfig—kpt_thresh,blur_kernel(DARK modulation, odd),box_pad,mean,std(RGB, applied after scaling to[0,1]).WholeBodyCrop—x1, y1, pad_left, pad_top, padded_w, padded_h: how a person box became the model canvas, and what inverts it.WholeBodyEstimator(engine, config=None, output_index=0)—.estimate(bgr, box, timeout_ms=0),.postprocess(crop),.config.wholebody_preprocess(bgr, x1, y1, x2, y2, in_w=192, in_h=256, config=None)→(input, crop). Takes BGR like the rest of this API and flips to the RGB the model wants. The box is widened bybox_pad, zero-padded to the model's 3:4 aspect and resized — not the mmpose center/scale affine.decode_wholebody(heatmaps, crop, config=None) -> list[Keypoint]— numpy path;heatmapsis a channel-first[K,H,W](or[1,K,H,W]) float array.
The model upscales one fixed tile; an arbitrary image is cut into overlapping tiles and cross-faded back together.
sr = bcdl.SuperResolver(engine) # scale and tile come from the model
big = sr.upscale(img) # BGR in, BGR out, sr.scale x larger
print(sr.last_tile_count, "tiles")SuperResConfig—overlap(input pixels; also the cross-fade width).SuperResolver(engine, config=None, input_index=0, output_index=0)—.upscale(bgr, timeout_ms=0),.scale,.tile,.last_tile_count.plan_tiles(width, height, tile_w, tile_h, overlap)→ tile origins. The last tile on each axis is flush with the far edge, so the real overlap there can exceed the requested one.tile_weight(i, len, ramp)— the cross-fade weight, always > 0 so the normalized blend needs no border special-casing.
Tile size is a deployment decision, not just a memory one: the compiled
.hbmscales with tile AREA (256x256 → 148 MB, 128x128 → 37 MB for the same network) at identical per-pixel throughput.
Repeatable keypoints with L2-normalized 64-d descriptors, plus mutual-NN matching. Only the convolutional trunk is on the BPU (~1.0 ms); the input normalization and all of the softmax / NMS / top-k / sampling run on the CPU.
ext = bcdl.FeatureExtractor(engine)
fa, fb = ext.extract(img_a), ext.extract(img_b)
for m in bcdl.match_features(fa, fb, min_cossim=0.82):
print(fa.xy[m.a], "->", fb.xy[m.b], m.score)XfeatConfig—detection_thresh,nms_kernel(odd),top_k.FeatureSet—keypoints: list[Feature],descriptorsas an(N, 64)array,xyas an(N, 2)array,dim,len().Feature—x, y, score.FeatureMatch—a, b, score.FeatureExtractor(engine, config=None, output_base=0)—.extract(bgr, timeout_ms=0),.postprocess(scale_x, scale_y),.config.xfeat_preprocess(bgr, in_w=640, in_h=480)→(input, scale_x, scale_y). Grayscale is the plain CHANNEL MEAN, not a luma weighting.decode_xfeat(feats, keypoints, heatmap, config=None, scale_x=1, scale_y=1)— numpy path; the three maps are[C,H,W](or[1,C,H,W]) float arrays with 64 / 65 / 1 channels.match_features(a, b, min_cossim=0.82)— mutual nearest neighbours. Cost isO(|a|*|b|*64): about 130 ms per pair at the defaulttop_kof 4096, 8 ms at 1024. Lowertop_kbefore anything else if matching dominates.
LTRB boxes plus a per-instance binary mask assembled from a prototype tensor.
cfg = bcdl.InstanceSegConfig(); cfg.compute_masks = True
seg = bcdl.InstanceSegmenter(engine, cfg)
for m in seg.detect([y_plane, uv_plane], lb, orig_w, orig_h):
print(m.class_id, m.score, m.mask.shape) # m.mask: (H, W) uint8 0/1InstanceSegConfig—conf_thresh,iou_thresh,max_dets,strides,proto_index,compute_masks(setFalseto skip mask assembly for speed).InstanceMask—x1, y1, x2, y2, score, class_id,mask_w, mask_h,mask((H,W)uint8 0/1; empty whencompute_masks=False).InstanceSegmenter(engine, config=None, output_base=0)—.detect([inputs], lb, orig_w, orig_h, timeout_ms=0),.postprocess(lb, orig_w, orig_h),.config.decode_instance_seg(cls, box, mc, proto, config, letterbox, orig_w, orig_h) -> list[InstanceMask]— numpy path;cls/box/mcper-stride[H,W,nc]/[H,W,4]/[H,W,np],proto[mH,mW,np].
LTRB plus an angle; rotated-IoU NMS.
cfg = bcdl.ObbConfig(); cfg.num_classes = 15
obb = bcdl.ObbDetector(engine, cfg)
for o in obb.detect([y_plane, uv_plane], lb):
r = o.rrect
print(o.class_id, o.score, r.cx, r.cy, r.w, r.h, r.angle) # angle in radObbConfig—num_classes,conf_thresh,iou_thresh,max_dets,strides,regularize,angle_offset_rad,angle_sign.RotatedBox—cx, cy, w, h, angle(radians).ObbDetection—rrect: RotatedBox,score,class_id.ObbDetector(engine, config=None, output_base=0)—.detect([inputs], lb, timeout_ms=0),.postprocess(lb),.config.decode_obb(cls, box, angle, config, letterbox) -> list[ObbDetection]— numpy path; per-stride[H,W,nc]/[H,W,4]/[H,W,1].rotated_iou(a_cx,a_cy,a_w,a_h,a_angle, b_cx,b_cy,b_w,b_h,b_angle) -> float.
Argmax over a logit tensor (or pass through a pre-argmaxed id tensor) to a per-pixel label map.
cfg = bcdl.SegConfig(); cfg.num_classes = 19
seg = bcdl.Segmenter(engine, cfg)
mask = seg.segment(x) # SegMask
labels = mask.labels # (H, W) int32 class ids
bgr = bcdl.seg_colorize(mask) # (H, W, 3) uint8 palette image
# numpy path:
mask = bcdl.decode_seg(logits.astype(np.float32), cfg)SegConfig—num_classes,channels_first(NCHW vs NHWC),argmaxed(setTrueif the model already emits ids).SegMask—width,height,num_classes,labels((H,W)int32).decode_seg(output, config) -> SegMask.seg_colorize(mask) -> np.ndarray—(H,W,3)uint8 fixed palette.Segmenter(engine, config=None, output_index=0)—.segment(model_input, timeout_ms=0),.postprocess(),.config.
Point cloud in, oriented 3-D boxes out (centre, extent, yaw in the ego frame, metres). A pillar detector is a pipeline, not one model, and the pieces stay separate because their costs are wildly different:
cfg = bcdl.VoxelConfig() # KITTI geometry by default
vox, npts, coords, stats = bcdl.voxelize_pillars(points_xyzi, cfg)
x = bcdl.pillar_point_features(vox, npts, coords, stats.num_pillars, cfg)
feats = vfe_engine.infer([x])[0] # BPU: (1, C, P, 1)
canvas = bcdl.scatter_pillars(feats.reshape(feats.shape[1], -1).T[:stats.num_pillars],
coords[:stats.num_pillars], stats.num_pillars, cfg)
head_engine.infer([canvas]) # BPU (backbone fused in)
boxes = bcdl.LidarDetector3d(head_engine).postprocess()Points must already be in the ego frame — nothing downstream can tell that a
cloud is still in sensor coordinates. And the caps are real: stats reports
points dropped out of range, points dropped because their pillar was full, and
occupied columns that had no slot left, so a sweep that overflows the model says
so instead of quietly going blind.
Two VFE forms exist and the difference is 179×. A VFE that takes the three
raw pillar tensors expands them into per-point features inside its own graph, and
pays 290 ms for a 95 MFLOP MLP — (pillars, points, 4) is a huge pillar axis
against a 4-wide channel axis, which is the worst shape a convolution engine can
be given. The same MLP re-exported as a 1×1 convolution over (1, 10, P, 32)
runs in 1.62 ms and takes pillar_point_features() instead. Feed whichever the
model asks for: a one-input VFE is the convolution form. Likewise a head whose
input is the scatter canvas has the backbone fused into it and needs no separate
backbone call. See docs/MODELS.md for both chains and how to get them.
VoxelConfig—min_x/max_x/min_y/max_y/min_z/max_z,pillar_x/pillar_y,max_points_per_pillar,max_pillars,num_features; read-onlygrid_x/grid_y.voxelize_pillars(points, config) -> (voxels, num_points, coords, Pillars)—pointsis(N, >=3)float32 x, y, z[, intensity…].pillar_point_features(voxels, num_points, coords, num_pillars, config) -> np.ndarray—(1, num_features+6, max_pillars, max_points): the raw point, its offset from the pillar's centroid, its offset from the pillar's centre. Unfilled point slots stay zero — they go through the MLP too, and the max over points would otherwise pick up whatever the bias maps them to.scatter_pillars(features, coords, num_pillars, config) -> np.ndarray—(1, C, grid_y, grid_x).Box3D—x,y,z,dx,dy,dz,heading,label,score.Anchor3dConfig— grid + per-classsizes/centers+rotations.generate_anchors_3d(config)returns the(N, 7)grid; it reproduces the published model's own anchor constants to 8e-06.Det3dConfig—score_thresh,nms_iou,max_boxes,apply_sigmoid,anchors.decode_det3d(cls, box, dir=None, anchors=..., config=...) -> list[Box3D]Passingdir=Noneleaves every heading correct only modulo pi.bev_iou(a, b)·nms3d(boxes, iou_thresh, max_boxes)— the metric this family's NMS is defined on is the BEV footprint, not the volume.LidarDetector3d(engine, config=None, cls_index=0, box_index=1, dir_index=2)—.postprocess(),.anchors.
"Is this part like the good ones." There is no class list and no training set of defects — the model is shown only normal examples of one product and reports how unlike them each patch looks.
Two consequences the API cannot hide: the model is per-product, so the
shipped .hbm is an example rather than a general detector (a new part means
retraining); and the threshold is a per-line decision calibrated against that
product's normal variation, not a library constant.
det = bcdl.AnomalyDetector(engine)
m = det.detect(bgr) # one (H, W, 3) uint8 BGR image
m.score # image-level verdict = max of the map
heat = m.data # (H, W) float32, at the model's input size
mask = bcdl.anomaly_mask(m, 0.5) # (H, W) uint8 0/255
# numpy path:
m = bcdl.decode_anomaly(grid.astype(np.float32), cfg)The image score is the map's maximum, not its mean: a part is as anomalous as its worst patch, and a mean would let one large clean region hide a small defect — the case the method exists to catch.
AnomalyConfig—out_width/out_height(0 ⇒ keep the raw grid; the detector defaults them to the model's input size),blur_kernel(odd, 0 disables),blur_sigma,threshold. The upsample and the Gaussian smoothing are part of the reference post-processing, not decoration.AnomalyMap—width,height,score,data((H,W)float32).decode_anomaly(output, config=AnomalyConfig()) -> AnomalyMap.anomaly_mask(map, threshold) -> np.ndarray—(H,W)uint8 0/255.anomaly_preprocess(bgr, in_w, in_h) -> np.ndarray— resize + BGR→RGB +/255+ ImageNet mean/std into(3,H,W).AnomalyDetector(engine, config=None, output_index=0, input_index=0)—.detect(bgr, timeout_ms=0),.postprocess(),.input_size.
Two frames in, a per-pixel displacement field out. Vectors are in source
pixels, +u right, +v down (the OpenCV convention), so the field drops straight
into cv2.remap once split into x/y maps.
est = bcdl.OpticalFlowEstimator(engine)
flow = est.estimate(frame0, frame1) # two (H, W, 3) uint8 BGR frames
uv = flow.data # (H, W, 2) float32, source pixels
bgr = bcdl.flow_colorize(flow) # Middlebury colour wheel
# numpy path:
flow = bcdl.decode_flow(out.astype(np.float32), cfg)
epe = bcdl.flow_endpoint_error(flow_a, flow_b)Preprocessing is inside the estimator on purpose: the reference implementation feeds BGR in 0-255 and divides by 255 inside the graph, so handing it RGB or pre-normalised pixels returns a plausible field rather than an error.
FlowConfig—scale_x,scale_y(model pixels → source pixels;estimate()sets them from the sizes it was given),channels_first([1,2,H,W]vs[1,H,W,2]).FlowField—width,height,data((H,W,2)float32 interleaved).decode_flow(output, config=FlowConfig()) -> FlowField.flow_colorize(flow, max_magnitude=0.0) -> np.ndarray—(H,W,3)uint8 BGR;0normalises by the field's own 99th-percentile magnitude so one fast object cannot flatten the rest of the frame.flow_endpoint_error(a, b) -> float— mean endpoint error in pixels, the metric to use when scoring a build against its float reference.flow_preprocess(bgr, in_w, in_h) -> np.ndarray— one frame to(3,H,W)float, BGR and 0-255 preserved.OpticalFlowEstimator(engine, config=None, output_index=0, input0_index=0, input1_index=1)—.estimate(frame0, frame1, timeout_ms=0),.postprocess(),.input_size.
The only head here that returns a decision rather than an observation: three forward cameras plus a lidar sweep in, one planned ego trajectory out, along with the perception heads it was trained jointly with. The multi-modal diffusion decoder lives inside the compiled graph — clustered trajectory anchors go in, one already-selected trajectory comes out.
All poses are ego frame, metres, x forward / y left, heading in radians CCW, and are the FUTURE at a fixed interval (8 poses over 4 s for the stock model).
anchors = np.load("kmeans_navsim_traj_20.npy").astype(np.float32) # (20, 8, 2)
planner = bcdl.DrivePlanner(engine, anchors)
camera = bcdl.stitch_drive_cameras(cam_left, cam_front, cam_right) # (3, H, W) RGB [0,1]
lidar = bcdl.lidar_bev_histogram(points_xyz) # (C, H, W)
status = bcdl.DriveStatus()
status.driving_command = 0 # route hint, one-hot slot
status.velocity_x = 5.0 # m/s
r = planner.plan(camera, lidar, status)
[(p.x, p.y, p.heading) for p in r.trajectory] # the plan
[(a.x, a.y, a.score) for a in r.agents] # detected vehicles
bgr = bcdl.seg_colorize(r.bev) # BEV semantic map (SegMask)The two sensor builders are geometric contracts, not conveniences: a transposed BEV grid or a status vector in the wrong order still infers and still returns a smooth-looking plan.
lidar_bev_histogram(points, config=LidarBevConfig()) -> np.ndarray—(N,>=3)ego-frame xyz →(C,H,W)occupancy in[0,1]. Grid axis 0 is x (forward), axis 1 is y.LidarBevConfig—min_x/max_x/min_y/max_y,pixels_per_meter,max_height,split_height,hist_max_per_pixel,use_ground_plane(false ⇒ one channel, points above the split only).stitch_drive_cameras(left, front, right, config=DriveCameraConfig()) -> np.ndarray— three(H,W,3)BGR images →(3,out_h,out_w)RGB float in[0,1].DriveCameraConfig—crop_top,crop_bottom,crop_side(left/right cameras only),out_width,out_height. The resize happens after concatenation, so pixels interpolate across the seams.DriveStatus—driving_command(one-hot slot),velocity_x,velocity_y,accel_x,accel_y.drive_status_vector(status, num_commands=4)packs it.DriveAnchorConfig— normalization (x_offset/x_span/y_offset/y_span), DDIM schedule (truncated_timestep,num_train_timesteps,beta_start,beta_end),noise_scale,seed.drive_anchor_noise(anchors, config)applies it;noise_scale=0zeroes the random term (reproducible comparisons) while keeping the schedule's scaling.DriveResult—trajectory(list ofDrivePose),agents(list ofDriveAgent:x,y,heading,length,width,score),bev(aSegMask, soseg_colorizeworks on it). The BEV map's axes are not screen axes: row index increases with FORWARD distance (row 0 is the ego end) and column index increases to the LEFT, so drawing it north-up needsbev.labels[::-1, ::-1]. Upstream documents neither, and a straight road is nearly symmetric under a 180° rotation — the wrong orientation looks entirely plausible.DriveConfig—agent_score_thresh,bev(aSegConfig).DriveIoConfig— which Engine tensor is which; defaults match the stock export. Set an output index to-1to skip that head.decode_drive(trajectory=None, agent_states=None, agent_labels=None, bev=None, config=DriveConfig()) -> DriveResult— the numpy path; any head may beNone.DrivePlanner(engine, anchors, config=None, anchor_config=None, io=None)—.plan(camera, lidar, status, timeout_ms=0),.postprocess(),.diffusion_input.
Single-channel depth / disparity to a float map, with colorization helpers.
cfg = bcdl.DepthConfig(); cfg.normalize = True
est = bcdl.DepthEstimator(engine, cfg)
dm = est.estimate(x) # DepthMap
arr = dm.data # (H, W) float32
bgr = bcdl.depth_colorize(dm) # turbo colormap (H, W, 3) uint8
g8 = bcdl.depth_to_gray8(dm) # (H, W) uint8
# numpy path:
dm = bcdl.decode_depth(out.astype(np.float32), cfg)DepthConfig—width,height,normalize(scale to [0,1]),clip_lo,clip_hi.DepthMap—width,height,vmin,vmax,data((H,W)float32).decode_depth(output, config) -> DepthMap.depth_colorize(dm) -> np.ndarray·depth_to_gray8(dm) -> np.ndarray.DepthEstimator(engine, config=None, output_index=0)—.estimate(model_input, timeout_ms=0),.postprocess(),.config.
Takes a depth map you already have — a stereo or ToF reading, noisy and full of
holes — plus the aligned RGB frame, and returns cleaned metric depth with a
per-pixel trust mask. This is the LingBot-Depth head; it does not estimate depth
from scratch the way DepthEstimator (monocular) or StereoPipeline do.
refiner = bcdl.DepthRefiner(engine) # inputs: image, depth_log
r = refiner.run(bgr, depth_m) # BGR (H,W,3) uint8 + (H,W) float32 metres
r.depth # (H, W) float32 metres, 0 where rejected
r.mask # (H, W) uint8, 1 = trusted
r.vmin, r.vmax # range over trusted pixels
k = bcdl.Intrinsics(fx, fy, cx, cy) # in the SENSOR image's pixels
k = bcdl.scale_intrinsics(k, bgr.shape[1], bgr.shape[0], r.width, r.height)
pts = bcdl.depth_to_pointcloud(r, k) # (H, W, 3) float32 metres, camera space
# numpy path — feed the graph yourself:
img = bcdl.preprocess_refine_image(bgr) # (1, 3, eh, ew) float32
dep = bcdl.preprocess_refine_depth(depth_m) # (1, 1, eh, ew) float32
r = bcdl.decode_refined_depth(depth_out, mask_logit_out)DepthRefineConfig—encoder_height,encoder_width(taken from the model at construction),min_valid_depth,mask_threshold,apply_mask.RefinedDepth—width,height,vmin,vmax,depth,mask.Intrinsics(fx, fy, cx, cy)·scale_intrinsics(k, from_w, from_h, to_w, to_h).preprocess_refine_image(bgr, config=None)·preprocess_refine_depth(depth_m, config=None)— the exact host preprocessing the model was calibrated with. The RGB downscale is area-averaged; a plain bilinear resize is measurably worse.decode_refined_depth(depth, mask_logit=None, config=None) -> RefinedDepth.depth_to_pointcloud(refined, intrinsics) -> np.ndarray.DepthRefiner(engine, config=None, image_input=0, depth_input=1, depth_output=0, mask_output=1)—.run(bgr, depth_m),.postprocess(),.config.
Depth goes into the graph as log depth with 0 meaning "no reading", which is upstream's own convention — a genuine 1 m reading and a missing pixel land on the same value, and the model leans on the RGB tokens to separate them. The deployed graph also keeps every depth token, where upstream drops tokens whose patch has no valid reading at all (a data-dependent sequence length cannot be compiled). On the reference scenes that costs 0.06% absolute relative error; those scenes are 87–100% valid, so very sparse input depth is outside what was measured.
Two rectified images to a disparity map, optionally metric depth and a validity
mask. Pixel normalization is fused into the .hbm; the C++ core does the fit +
BGR→RGB + F32 NCHW pack. The fit mode must match how the model was
calibrated.
cfg = bcdl.StereoConfig()
cfg.fit = bcdl.StereoFit.Crop # or Resize — MUST match calibration
cfg.fx, cfg.baseline = 700.0, 0.12 # enable metric depth (z = fx*baseline/disp)
cfg.valid_mask = True
pipe = bcdl.StereoPipeline(engine, cfg)
res = pipe.process(left_bgr, right_bgr)
disp = res.disparity.data # (H, W) float32 disparity (a DepthMap)
depth = res.depth # (H, W) float32 metric depth, or shape (0,) if off
valid = res.valid # (H, W) uint8 mask, or shape (0,) if offStereoFitenum —Resize,Crop.StereoConfig—input_w,input_h,fit,to_rgb,left_index,right_index,output_index,fx,baseline,valid_mask,disp_min,max_disp,left_margin,lr_check,lr_thresh.StereoResult—disparity: DepthMap,depth(numpy, empty when nofx/baseline),valid(numpy uint8, empty whenvalid_mask=False).StereoPipeline(engine, config=None)—.process(left, right) -> StereoResult,.input_w,.input_h.- numpy helpers:
pack_stereo_input(bgr, out_h, out_w, fit=StereoFit.Resize, to_rgb=True),disparity_to_depth(disp, fx, baseline),stereo_valid_mask(disp, disp_min=0.0, max_disp=192.0, left_margin=..., lr_check=False, lr_thresh=...).
Full PP-OCR three-stage pipeline (PP-OCRv6 by default, v5 kept as a fallback),
each stage usable on its own. Pure decoders (decode_dbnet, decode_cls_dir,
decode_ctc, load_char_dict) need no model.
chars = bcdl.load_char_dict("ppocr_dict.txt")
det = bcdl.DbTextDetector(engine_det) # DBNet detect
boxes = det.postprocess(lb) # list[TextBox], 4-point rotated
cls = bcdl.TextAngleClassifier(engine_cls) # 0/180 direction
rec = bcdl.TextRecognizer(engine_rec, "ppocr_dict.txt") # CRNN/CTC
# (crop each box, run cls then rec on the crop's engine; see examples/ocr_demo)
# numpy decoders:
boxes = bcdl.decode_dbnet(prob.astype(np.float32), bcdl.DbConfig(), lb)
dir_ = bcdl.decode_cls_dir(logits.astype(np.float32), thresh=0.9)
text = bcdl.decode_ctc(logits.astype(np.float32), chars)DbConfig—bin_thresh,box_thresh,unclip_ratio,min_size,connectivity.TextBox—x1, y1, x2, y2,score,points((4,2)float, clockwise, original pixels).RecResult—text,score.ClsDirResult—label,score,flip180.load_char_dict(path) -> list[str]— one token per line.decode_dbnet(prob, config, letterbox) -> list[TextBox]— connected-component + unclip on a(H,W)probability map.decode_cls_dir(logits, thresh=0.9) -> ClsDirResult.decode_ctc(logits, dict) -> RecResult— CTC greedy decode of a(T,C)array.- Engine-bound:
DbTextDetector(engine, config=None, output_index=0),TextAngleClassifier(engine, thresh=0.9, output_index=0),TextRecognizer(engine, dict_path, output_index=0)— each has.postprocess(...).
See examples/ocr_demo.cc for the full det→cls→rec
wiring (crop ordering and the dict gotchas are handled there).
Model-free: feed each frame's detections, get stable track ids back.
tracker = bcdl.ByteTracker(bcdl.ByteTrackConfig())
for frame in stream:
dets = detector.detect(...) # list[Detection]
for t in tracker.update(dets):
print(t.track_id, t.class_id, t.x1, t.y1, t.x2, t.y2)ByteTrackConfig—track_thresh,high_thresh,match_thresh,track_buffer,frame_rate; appearance:proximity_thresh,appearance_thresh,ema_alpha;boost(BoostConfig).Track—track_id,x1, y1, x2, y2,score,class_id;__repr__.ByteTracker(config=None)—.update(detections) -> list[Track],.update(detections, embeddings) -> list[Track],.apply_camera_motion(affine),.reset(),.config.
Passing one embedding per detection switches on appearance association: the cost
becomes min(IoU distance, gated cosine distance), so appearance can only
rescue a match geometry was about to miss — it can never break one geometry
already had. An empty entry means "no appearance for this detection", which
is how you skip the ReID model on cheap crops.
reid = bcdl.ReIDExtractor(bcdl.Engine("osnet.hbm"))
tracker = bcdl.ByteTracker(bcdl.ByteTrackConfig())
for frame in stream:
dets = detector.detect(...)
embs = reid.embed_detections(frame, dets, min_score=0.5) # parallel to dets
for t in tracker.update(dets, embs):
print(t.track_id)ReIDExtractor(engine, config=None, output_index=0)—.dim,.embed(model_input),.embed_crop(crop_bgr),.embed_detections(frame_bgr, detections, min_score=0.0).reid_preprocess(crop_bgr, width=128, height=256, config=None)andreid_crop_preprocess(bgr, x1, y1, x2, y2, in_w, in_h, config)— a squashing resize (not a letterbox: these models are trained on squashed crops), BGR→RGB, ImageNet normalization, NCHW float32.ReidConfig—mean,std.
The model runs once per crop, so on a crowded frame it, not the detector,
sets the frame time. Note that int8 PTQ is not adequate for OSNet — see the
0.4.0 entry in CHANGELOG.md.
apply_camera_motion(affine) warps every tracklet by a 2x3 transform mapping
the previous frame onto the current one, before the next update(). Position
and size are warped; velocity is not, because velocity describes the target in
the world and a one-frame camera jolt is not the target accelerating. Skip the
call entirely for a static camera rather than passing an identity.
M, _ = cv2.estimateAffinePartial2D(prev_pts, cur_pts, method=cv2.RANSAC)
tracker.apply_camera_motion(np.ascontiguousarray(M, np.float32))BoostTrack++ additions, all off by default — each is independently switchable so its contribution can be measured rather than assumed.
rich_similarity— add Mahalanobis and shape-agreement terms to the cost;lambda_iou,lambda_mhd,lambda_shape,min_iou.soft_biou— grow both boxes by1 - tracklet confidence, so a tracklet that has been coasting searches a wider area.boost_detections— raise the scores of detections the tracklets vouch for, before the high/low split;dlo_alpha,vt_start,vt_end,vt_steps,duo,duo_iou. Situational: it recovers real detections when the detector is limited by misses, and manufactures false tracks when it is not.
Detect-and-track in one call; all preprocessing in C++. Needs an NV12-input
YOLO .hbm.
pipe = bcdl.TrackingPipeline(engine) # det_config, track_config optional
for frame in stream: # frame: HxWx3 uint8 BGR
for t in pipe.process(frame):
print(t.track_id, t.class_id, t.score)
print(pipe.last_detections) # pre-association dets of last framePipelineConfig—input_w,input_h,detect(DetectConfig),output_index,pad_value,head(DetectHead),ltrb_strides.DetectHeadenum —Auto,SingleTensor,YoloLtrb.TrackingPipeline(engine, det_config=None, track_config=None, reid_engine=None, reid_config=None)—.process(bgr) -> list[Track],.last_detections,.has_reid,.last_embed_count,.reset(). Passingreid_engineadds appearance to the native path; the crop size is read from that model.TrackingReidConfig—min_score,max_crops,crop(ReidConfig). Both are cost knobs: the ReID model runs once per qualifying crop, andlast_embed_countreports how many that was.
Streaming detection with CPU preprocessing of later frames overlapped against BPU
infer+decode of earlier ones. submit() / next() block (backpressure / wait)
but release the GIL. Results return in submission order. Needs an NV12-input
YOLO .hbm.
cfg = bcdl.PipelineConfig(); cfg.detect.num_classes = 80
pipe = bcdl.AsyncDetectionPipeline(engine, cfg, depth=3)
for i, frame in enumerate(stream): # frame: HxWx3 uint8 BGR
pipe.submit(frame) # blocks if full
if i >= 3:
for d in pipe.next(): # in submission order
...
pipe.finish()
while (dets := pipe.next()) is not None:
... # drain in-flight framesAsyncDetectionPipeline(engine, config=None, depth=3):submit(bgr) -> bool— enqueue (bytes copied; array reusable immediately). Blocks while full; returnsFalseafterfinish().next() -> list[Detection] | None— pop next result;Noneonce finished and drained.finish()— signal end of stream (idempotent; also runs on GC).head -> DetectHead— resolved decoder family.profile() -> StageProfile— per-stage service time (preproc_ms/infer_ms/postproc_mstotals +*_per_frame()); the slowest stage bounds throughput. Read afterfinish()+ drain.
The whole compressed-video → detections path in C++ (decode ‖ nv12→bgr ‖ preproc ‖ infer+NMS, four overlapped stages). Python only pumps bytes — feed Annex-B
chunks (e.g. an ffmpeg -c copy RTSP/mp4 stream) and read detections; all
decode/convert/detect threads run in C++ with the GIL released. This is what lets
a thin Python driver hit the C++ decode-bound ceiling (~441 FPS @1080p H.264)
instead of the ~81 FPS a Python-orchestrated decode loop is capped at.
cfg = bcdl.PipelineConfig(); cfg.detect.num_classes = 80
pipe = bcdl.AsyncVideoDetectionPipeline(engine, cfg, bcdl.VideoType.H264, depth=4)
while chunk := ffmpeg.stdout.read(65536): # ffmpeg -c copy (no software decode)
pipe.submit(chunk) # AUs segmented + VPU-decoded in C++
while (dets := pipe.next_nowait()) is not None:
... # drain what's ready
pipe.finish()
while (dets := pipe.next()) is not None: # blocking drain of last frames
...AsyncVideoDetectionPipeline(engine, config=None, codec=VideoType.H264, depth=4):submit(bytes) -> bool— feed Annex-B bytes; blocks on backpressure. An MP4 is not Annex-B (AVCC length prefixes, SPS/PPS in theavcCbox), so feeding one yields zero frames. Demux it first — container only, pixels untouched, VPU still the only decoder:ffmpeg -i in.mp4 -c:v copy -bsf:v h264_mp4toannexb -f h264 -. Seeload_annexb()inexamples/video_det_demo.py.next() -> list[Detection] | None— blocking pop in decode order.next_nowait() -> list[Detection] | None— non-blocking pop.finish(),profile() -> StageProfile(incl.decode_ms,cvt_ms).
- Video decode handles H.264 and H.265 (a hierarchical-GOP HEVC stream decodes its base temporal layer only — see the VideoDecoder note above).
- See
examples/rtsp_det_demo.py— a thin RTSP driver built on this (ffmpeg pipe →submit→next_nowait).
examples/video_det_demo.py is the end-to-end
Python demo: VPU decode → AsyncDetectionPipeline (BPU) → draw → VPU
encode → .mp4. It reads raw .h264/.h265 or .mp4/.mov (MP4 is demuxed to
Annex-B with ffmpeg -c copy; the VPU still does the actual decode), and muxes
the VPU's elementary stream into .mp4 with ffmpeg -c copy (container only, no
re-encode). Prints the per-stage profile() distribution.
python examples/video_det_demo.py det.hbm in.mp4 out.mp4 # mp4 -> mp4
python examples/video_det_demo.py det.hbm in.h264 out.mp4 300 4 # [max_frames] [depth]Note: decode and encode share the single VPU core, so the full round-trip runs slower (~104 FPS @1080p yolo26n, encode-bound at 3.56 ms/frame) than the decode-only detect path (~441 FPS).
The unified shared-memory image plus the JPU (JPEG) and VPU (H.264/H.265) hardware codecs. These allocate cached device buffers and drive the media units, so they only run on the board.
# JPEG (JPU) — one-shot helpers:
jpg = bcdl.jpeg_encode(bgr, quality=80) # bytes
img = bcdl.jpeg_decode(jpg) # VpImage (NV12)
planes = img.to_numpy() # see VpImage.to_numpy below
# Reuse a decoder for a stream (re-creating one per call adds ~5 ms JPU setup):
dec = bcdl.JpegDecoder()
for blob in blobs:
vp = dec.decode(blob)
# H.264 / H.265 (VPU):
ec = bcdl.VideoEncConfig(); ec.type = bcdl.VideoType.H264
ec.width, ec.height, ec.bitrate_kbps, ec.framerate = 1280, 720, 4000, 30
enc = bcdl.VideoEncoder(ec)
chunk = enc.encode(vp_image) # bytes (may be empty if buffered)
dc = bcdl.VideoDecConfig(); dc.type = bcdl.VideoType.H265 # or .H264
dec = bcdl.VideoDecoder(dc)
frame = dec.decode(nal_bytes) # VpImage or None while buffering
# Decoupled feed/drain (handles H.265 reorder): feed arbitrary bytes, drain in
# display order, flush the reorder tail at end-of-stream.
dec.feed(chunk_bytes) # queue bytes (no AU split needed)
while (f := dec.receive(0)) is not None: ... # non-blocking drain (0 = no wait)
while (f := dec.flush()) is not None: ... # after last feed: reorder tailH.265 note: the decoder uses the reorder-correct
media_codecAPI, so H.265 works (an older per-AU model timed out on it). A hierarchical-GOP HEVC stream (Hikvision SVC / "H.265+"/smart-codec) decodes base temporal layer only (~1/4 of frames) — an SoC-SDK limitation; disable the camera's SVC mode for full rate. H.264 and non-hierarchical HEVC decode fully.
ImageFormatenum —Y,NV12,RGB,BGR.VpImage(width, height, format)—width,height,format,valid;.to_numpy()copies out honoring the device row stride:BGR/RGB -> (H,W,3),Y -> (H,W),NV12 -> flat (W*H*3//2)(Y plane then interleaved UV).vp_image_from_bgr(bgr) -> VpImage·vp_image_from_nv12(nv12, width, height) -> VpImage.JpegEncoder(width, height, quality=50, format=NV12)—.encode(VpImage) -> bytes(width must align to 16, height to 8).JpegDecoder(out_format=NV12)—.decode(bytes) -> VpImage. Reuse one instance across a stream; constructing per call costs ~5 ms.jpeg_encode(bgr, quality=50) -> bytes·jpeg_decode(data) -> VpImage— one-shot convenience wrappers (need OpenCV + even dims).VideoTypeenum —H264,H265.VideoEncConfig—type,width,height,bitrate_kbps,framerate,intra_period,format.VideoEncoder(config)—.encode(frame: VpImage) -> bytes(empty when the encoder buffered the frame);.type/.width/.height/.format.VideoDecConfig—type,format,in_buf_size.VideoDecoder(config)—.decode(data) -> VpImage | None(Nonewhile buffering reference frames);.type/.format.
The C++ surface mirrors the above; configs and result structs share field names
(camelCase). Start from include/bcdl/bcdl.h and the
per-area headers:
| area | header |
|---|---|
| core (SysMem, Task, Status, MemPool) | core/ |
| backend (Engine, output readers) | backend/engine.h |
| preprocessing (letterbox, VpImage) | preproc/ |
| media (JPEG/video codecs) | media/ |
| tasks (det/cls/pose/seg/obb/depth/ocr) | tasks/ |
| tracking | tracks/byte_tracker.h |
| pipelines | pipeline/ |
Errors are reported via BCDL_CHECK(...) raising bcdl::Error. See
examples/ for runnable programs.