Skip to content
dataglifLet’s talk
Field Notes

The details make the dataset.

Practical observations from annotation, custom tooling, and quality work. A closer look at how we turn complex data into usable structure.

Computer Vision

Combining expert annotation with SAM models.

Driving sequences require more than a good annotation on a single frame. Identities, boundaries, and motion must remain coherent as the scene changes.

GPU-accelerated inference and temporal tracking help generate consistent masks. Human annotators review and refine the results, preserving lane boundaries, vehicle trajectories, occlusions, and merge behavior.

Frame-level reviewTemporal trackingHuman correction
roundabout.mp4 · lanes, cuboids, tracking

Model-drafted, human-refinedFrame-level review

Custom Tooling

Rotation-aware cuboids for changing perspectives.

Driver-view footage constantly changes a vehicle’s angle and perspective. Default annotation controls do not always make those rotations intuitive.

Our rotation-aware cuboid tooling extends CVAT so annotators can align 3D labels more naturally and maintain consistency across frames.

CVAT extension3D alignmentMultiview consistency
cuboid-rotation.mp4 · CVAT extension
Document Intelligence

Layout is the context layer.

OCR can find the words. Structure reveals how they relate. Sections, tables, lists, and hierarchy carry context in procurement packages, job postings, and legal documents.

We build and validate that structural layer first, creating a clearer foundation for extracting skills, clauses, dates, and contract terms.

Open the document demo
Annotation Intelligence

Turn disagreement into better guidance.

Expert review and AI-assisted analysis help identify annotation disagreements, surface ambiguous guidelines, and show where the team needs more support.

The outcome is practical: targeted coaching, a clearer rubric, and more consistent annotations, with human judgment at the center.

Open the calibration demo

About Meridian Autonomy

Meridian Autonomy builds perception systems for autonomous heavy equipment operating across construction and mining sites nationwide. The company partners with leading equipment manufacturers to bring machine-learning-driven safety systems to hazardous work environments, and has annotated more than 40 million frames of multi-camera sensor data to date. Meridian Autonomy is expanding its annotation program to support a new generation of LiDAR-fused perception models, and is contracting the review work through Meridian Data Partners.

About the Opportunity

Meridian's annotation network is composed of in-house reviewers and contract specialists working directly inside CVAT and our internal QA tooling. With calibrated rubrics, a supportive review pipeline, and regular coaching from senior annotators, being part of the network is a chance to do meticulous, detail-oriented work that directly shapes how a model performs in the field.

As a contract Senior Annotation Specialist, you will provide frame-level cuboid annotation. This role emphasizes core competencies such as spatial reasoning, edge-case judgment, rubric adherence, and cross-frame consistency to ensure the highest standard of dataset quality.

Responsibilities:

- Annotate rotation-aware 3D cuboids across synchronized multi-camera frames

- Flag occlusion, truncation, and sensor artifacts according to the labeling rubric

- Review peer annotations and reconcile disagreements during weekly calibration sessions

- Experience and commitment to providing accurate, consistent spatial annotation that follows internal quality guidelines

Requirements:

- Graduate from an accredited program in Computer Science, Robotics, or a related field; OR equivalent hands-on annotation experience with preference for a Certified Data Annotation Specialist credential

- Reliable access to a dual-monitor annotation workstation

- A minimum of 1 year of annotation experience

- Strong preference for annotators with an active CVAT proficiency certification

- Must be available for weekly calibration reviews

The Bigger Picture

Why better data matters.

Growing demand brings the quality and usefulness of training data into focus. This published scenario offers context for the work: clearer structure, careful annotation, and consistent human review.

Training data availability

Data Exhaustion Projection: Training Data vs. Available Stock

  • Actual Models
  • Historical Trend
  • Compute-Optimal Projection
  • Data Bounds (9-27T tokens)
  • 2025 Baseline
Data Exhaustion Projection: Training Data vs. Available StockReconstructed model dataset sizes and a published scenario, using a 2025 baseline and a 9–27 trillion token stock range. The source projection reaches its upper bound around 2027.4. It does not establish a universal date when training data runs out.GPT-3Llama 2Llama 3DeepSeek-V32025 BaselineAvailable datastock range · 9–27T051015202530202020222024202620282030YearDataset Size (Trillion Tokens)Data Exhaustion Projection: Training Data vs. Available StockReconstructed model dataset sizes and a published scenario, using a 2025 baseline and a 9–27 trillion token stock range. The source projection reaches its upper bound around 2027.4. It does not establish a universal date when training data runs out.GPT-3Llama 2Llama 3DeepSeek-V32025 Baseline051015202530202020242028YearDataset Size (Trillion Tokens)
Published scenario · 2025 baseline. Values are approximate, reconstructed from the source figure; anonymous points remain unlabeled. This projection does not establish a universal date when training data runs out. Source: Patro & Agneeswaran, LLMOrbit (2026).
Model dataset sizes reconstructed from the source figure
ModelYearTrillion tokens
Unlabeled source point20190.1
GPT-320200.35
Unlabeled source point20211.4
Unlabeled source point20221.4
Unlabeled source point20231.4
Llama 220232
Unlabeled source point202312
Unlabeled source point202313
Unlabeled source point202410
Unlabeled source point202412
Llama 3202415
DeepSeek-V3202514.8
Build What Comes Next

The next breakthrough starts with better data.

Show us the data, the edge cases, or the gap. We’ll help shape the workflow that moves you forward.

Talk about your project