Transitioning from Raw Images to Structured Labeled Data in Loon Conservation

A weekly update on the Loon Conservation AI project covering the completion of CVAT image annotation and why the team is holding off on training a detection model until the dataset grows.

2:28 video3 min readWatch on YouTube

Swara Joshi gives a weekly update on the Loon Conservation AI project, where two tracks, the annotated dataset and the researcher-facing prototype, are converging toward the same system. The headline milestone is finishing CVAT annotation on every image currently available, which moves the project from raw footage to a structured, labeled dataset.

Finishing annotation in CVAT

Joshi and a collaborator, Nikhil, completed labeling every image the team currently has access to in CVAT, the annotation tool used to mark loons in the footage. The work meant handling real inconsistencies across the images: differences in object size, visibility, image quality, framing, lighting, reflections, and environmental conditions. Each of these has to be treated consistently, because consistency across the dataset matters more as the dataset grows. The stated pipeline is simple: raw images go through annotation, become a labeled dataset, and eventually feed model training and detection. Finishing this round of annotation is what pushes the project past the raw-image stage into having real structured data, which is the prerequisite for any training work.

Why the team isn't training a model yet

Finishing annotation naturally raises the question of whether it's time to start training a detection model. Joshi's answer is not yet. The dataset is still relatively small, and the plan is to keep annotating and building before attempting any training or evaluation. The reasoning is direct: a weak model trained too early would tell the team less than having no model at all. Rather than rush to a first training run, the strategy is to wait until there is enough loon-specific data to make that first attempt meaningful.

Prototyping the researcher-facing interface

In parallel with the annotation work, Joshi kept refining a prototype aimed at the researchers who will eventually use the tool. The prototype isn't wired to a real computer vision model yet, but it simulates how footage uploads, metadata, detections, observations, and verification will connect once the system is live. Even without a working model behind it, the prototype is already useful for identifying where the researcher experience needs to change before the real detection pipeline is built.

What comes next

The next phase depends on two things happening together: more footage being sourced and gathered, and feedback from the National Loon Center and Nikhil shaping the front end. Once there is enough loon-specific data, the plan calls for experimenting with detection and counting. Joshi frames the current state honestly: there's no trained model yet, and the dataset is still relatively small, but completing annotation on every currently available image is a genuine milestone that moves the project from raw images to structured, labeled data.

Key takeaways

  • Every currently available image has been fully annotated in CVAT, completing the shift from raw images to a structured labeled dataset.
  • Consistency in handling variation, size, visibility, lighting, framing, reflections, is treated as essential groundwork before training.
  • The team is deliberately not training a detection model yet because the dataset is still too small to produce a useful result.
  • A researcher-facing prototype is being refined in parallel, simulating the upload, detection, and verification workflow even without a live model.
  • Progress depends on sourcing more footage and incorporating feedback from the National Loon Center before moving to detection and counting experiments.

Who this is for

This update is for anyone following the Loon Conservation AI project, or for conservation and citizen-science teams building their own image-based monitoring tools who want a concrete example of sequencing annotation, prototyping, and training decisions.

Chapters

  1. 0:00Weekly Overview: Front-End and Data Tracks Converge
  2. 0:35Dataset Milestone: Wrapping up Image Annotation in CVAT
  3. 1:10The Discipline of Consistency: Documenting Decisions Across Environmental Changes
  4. 1:45Prototyping the Researcher UX: Simulating Uploads and Verification Loop
  5. 2:20The Strategic Wait: Why Small Datasets Shouldn't Be Trained Too Early
Full transcript(auto-generated, with timestamps)

Weekly Overview: Front-End and Data Tracks Converge

[0:00]Hi, this is Swara Joshi. And this week on the Loon Conservation AI project, two things came together. Nikhil and I finished annotating every image currently available to us in CVAT, and I kept refining the researcher-facing prototype so its workflow matches the real tool we're building toward. This week had a real through-line: finish annotating every image we currently have, and keep shaping the researcher-facing prototype so it matches the tool we're actually building. Two separate tracks, front end and data set, both moved from planning into structured hands-on work, and both are converging toward the same eventual system. Completing the annotations meant

Dataset Milestone: Wrapping up Image Annotation in CVAT

[0:36]Confronting real practical challenges: differences in object size, visibility, image quality, framing, lighting, reflections, and environmental conditions. Every one of those has to be handled consistently, because consistency across the data set only gets more important as it grows. The process is simple to state: raw images go through annotation, become a labeled data set, and eventually feed model training and detection. Completing this week's annotation work is the milestone that moves us from having raw images to having real, structured, labeled data, the necessary foundation before any model training can

The Discipline of Consistency: Documenting Decisions Across Environmental Changes

[1:11]Begin. One real question follows from finishing this passive annotation: is it time to start training a detection model? Not yet. The data set is still relatively small, so the plan is to keep annotating and building before training and evaluating anything. A weak model trained too early would tell us less than no model at all. On the interface side, the real work this week was thinking through how footage uploads, metadata, detections, observations, and verification will eventually connect. The prototype still isn't wired to a real computer vision model, but it's already useful for spotting where the researcher experience needs to change.

Prototyping the Researcher UX: Simulating Uploads and Verification Loop

[1:46]So, the honest state: we don't have a trained model, and the data set is still relatively small. But finishing the annotation of every currently available image is a real milestone. It moves this project from raw images to structured labeled data with the prototype's workflow shaped to match. Next, I'll keep building the data set as more footage becomes available and keep refining the front end based on feedback from the National Loon Center and Nikhil. Once there's enough loon specific data, the next real step is experimenting with detection and counting. If you want to help, here's one question. What's the clearest sign our data set is finally big and varied

The Strategic Wait: Why Small Datasets Shouldn't Be Trained Too Early

[2:21]Enough to start training on? That's the update for this week. Thank you for watching and I'll see you in the next one.

More from Madison

Humanitarians AI Lyrical Literacy Project