COMPUTER VISION

Image & Video Annotation Services

Pixel-perfect annotations for object detection, segmentation, tracking, pose estimation, and 3D point cloud labeling — with measurable IoU thresholds, per-class accuracy tracking, and multi-tier QA at every step. From single-class bounding boxes to complex multi-sensor fusion annotation.

5M+

Images Annotated

IoU ≥ 0.92

Avg Quality Score

50K+

Frames/Week Capacity

L1→L2→L3

QA Pipeline

< 2px

Boundary Precision

25+

Automated QA Rules

Capabilities

Six Core Annotation Capabilities

Bounding Box Annotation

2D axis-aligned bounding boxes, oriented (rotated) bounding boxes, and 3D cuboid annotations for object detection across all domains. Configurable overlap policies, occlusion flags, truncation percentage, and difficulty scoring per annotation.

Semantic & Instance Segmentation

Pixel-perfect semantic segmentation, instance segmentation with unique object IDs, and panoptic segmentation combining both. Support for 100+ class taxonomies with polygon, brush, superpixel, and SAM-assisted annotation tools.

Keypoint & Pose Estimation

Configurable skeleton definitions for human pose (17–133 keypoints), hand tracking (21 points), facial landmarks (68–478 points), and custom articulated objects. Each keypoint includes visibility flags (visible, occluded, out-of-frame) and confidence indicators.

Video Object Tracking

Multi-object tracking (MOT) with persistent IDs maintained through occlusions, re-entries, and camera transitions. Keyframe annotation with linear/spline interpolation and manual correction. Support for single-object tracking (SOT), MOT, and multi-camera cross-view tracking.

Image Classification & Tagging

Single-label, multi-label, and hierarchical classification with configurable confidence thresholds. Support for fine-grained recognition (breed identification, species classification), quality assessment (defect grading), and content moderation across millions of images.

3D Point Cloud & LiDAR Annotation

3D bounding cuboid annotation in LiDAR point clouds with heading angle, velocity estimation, and multi-frame tracking. Semantic segmentation of point clouds, lane/road boundary marking, and sensor fusion annotation linking LiDAR to camera imagery.

Autonomous Driving

LiDAR-camera fusion, 3D cuboid tracking, lane detection, traffic sign/light recognition, pedestrian tracking, and edge-case scenario annotation across ODD (Operational Design Domain) coverage matrices.

Medical Imaging

DICOM annotation for CT, MRI, X-ray, histopathology WSI, and ophthalmology. Organ segmentation, lesion classification, landmark detection, and measurement by board-certified radiologists and pathologists.

Retail Intelligence

Product recognition, shelf compliance analysis, planogram verification, visual search annotation, customer behavior tracking, and inventory management for retail AI systems.

Manufacturing QC

Surface defect detection, assembly verification, weld inspection, dimensional compliance, and quality grading for industrial quality control on high-speed production lines.

Surveillance & Security

Person detection, action recognition, anomaly detection, crowd analysis, license plate recognition, and multi-camera tracking with privacy-compliant annotation workflows.

Agriculture & Earth Observation

Drone and satellite imagery annotation for crop health monitoring, pest detection, land use classification, yield estimation, and infrastructure monitoring across growing seasons.

Gold Set Calibration

IoU threshold validation against expert-labeled gold sets. Annotators must achieve 0.85+ IoU on the gold set before touching production data. Gold sets refreshed 10% monthly to prevent memorization.

Multi-Reviewer Pipeline (L1→L2→L3)

L1 annotators produce initial labels. L2 reviewers audit 100% of output (not sampling). L3 adjudicators resolve disagreements and edge cases. Complex annotations always pass through at least two sets of eyes.

Per-Class Metrics Tracking

We track accuracy, precision, recall, and IoU per class — not just aggregate metrics. Rare but critical classes (pedestrians, small objects, defects) get extra QA attention and dedicated review queues.

Automated Consistency Checks

Rule-based validation catches common errors: overlapping bounding boxes, missing labels, impossible polygon shapes, label-class mismatches, and boundary violations. Errors flagged before human review.

Drift Detection & Alerts

Statistical monitoring across batches detects quality drift before it impacts your model. Batch-over-batch IoU, accuracy, and error-type distributions are compared. Automatic alerts trigger recalibration when drift exceeds ±2%.

Edge Case Libraries

Growing libraries of ambiguous and edge-case examples with documented resolution decisions. Used for annotator training, guideline refinement, and quality audit. Every edge case becomes a reusable training asset.

Comparison

UTL CV Annotation vs. Typical Providers

CapabilityUTL Data EngineTypical Provider
Gold set calibration with monthly refresh ✓ Initial only
100% L2 review (not sampling) ✓ 5–10% sampling
Per-class IoU/precision/recall tracking ✓ Aggregate only
3D cuboid + multi-sensor fusion ✓ 2D only
Automated consistency validation (25+ rules) ✓ Basic checks
Drift detection with auto-alerts ✓
Edge case libraries with decision docs ✓
SAM-assisted pre-annotation ✓
Domain-trained annotators (20+ hr onboarding) ✓ 2–4 hr training
Sub-pixel polygon accuracy (< 2px) ✓ 5–10px typical
““UTL reduced our annotation rework by over 50%. Their gold set calibration and per-class IoU tracking caught quality issues that our previous vendor missed entirely. The 100% L2 review coverage is what makes the difference — no more sampling-based QA surprises.””

ML Engineering Lead

Series B Autonomous Vehicle Company

FAQs

Computer Vision Questions

What IoU thresholds do you target?

Default targets: IoU ≥ 0.90 for bounding boxes, mIoU ≥ 0.88 for segmentation, OKS ≥ 0.85 for keypoints, 3D IoU ≥ 0.70 for cuboids. For precision-critical projects (medical, AV), we configure higher thresholds (≥ 0.95 for 2D, ≥ 0.92 for segmentation). All thresholds are agreed during scoping and validated against gold sets.

Can you handle 3D point cloud and multi-sensor fusion?

Yes. We annotate LiDAR point clouds with 3D cuboids, point-level segmentation, and trajectory tracking. Multi-sensor fusion links annotations across LiDAR, camera, radar, and IMU with ≤ 10ms timestamp synchronization. Our teams are trained on common AV sensor configurations.

What's your capacity for video annotation?

Steady-state: 500–2K video-minutes/week per project with MOTA ≥ 0.90 tracking accuracy. For burst projects, we scale to 5K+ video-minutes/week with parallel annotator teams. Multi-object tracking, re-identification, and temporal action segmentation are all supported.

How do you handle occlusion and edge cases?

We build edge-case libraries during guideline creation (100+ documented cases per mature project) and maintain a structured decision taxonomy. Occluded objects receive visibility flags (fully visible, partially occluded, heavily occluded) and truncation percentages. New edge cases discovered during production are documented and added to the library with resolution decisions.

Do you support SAM-assisted annotation?

Yes. We use Segment Anything Model (SAM) for pre-annotation to accelerate segmentation tasks. Human annotators review, correct, and validate all SAM-generated masks. This hybrid approach typically delivers 2–3× throughput improvement while maintaining our quality standards.

What's the typical engagement timeline?

Scoping + guideline design: 3–5 days. Team assembly + calibration: 5–7 days. Pilot (1K–5K samples): 5–10 days. First labeled batch by Day 20. Full production velocity by Day 25. We maintain pre-qualified teams across major domains for faster ramp-up.

Need Pixel-Perfect Annotations?

Let's discuss your computer vision data pipeline — from task design to quality-assured delivery. We'll scope a pilot within 48 hours.