Skip to content
MoeLink — Field Notes

From 14,000 Labeled Frames to a Monitored Edge Deployment in Nine Days: An FGCV Post-Mortem

A nine-day post-mortem: 14,000 labeled frames, a 120ms latency budget, and the data audit that mattered more than any model architecture.

A reader we'll call Dana runs computer vision for a mid-size logistics operator — the kind of company that never makes the trade press but quietly moves a few hundred thousand parcels a week. In early spring she wrote to us with a problem we hear constantly: her team had a detection model that worked beautifully in a notebook and a pilot deadline that was already slipping. Nine days later, the model was running on a dozen dock cameras, and Dana had a monitoring dashboard she trusted. This is how that happened, reconstructed from her notes and a follow-up call.

The project was mundane on purpose: flag damaged cartons on a conveyor before they reached a palletizer. The dataset existed — roughly 14,000 labeled frames, accumulated over two years by an annotation vendor and two in-house contractors. What did not exist was a repeatable path from that pile of images to something a night-shift supervisor could rely on. Dana's previous attempt had died in what she called "the integration swamp": a training script in one repo, an evaluation notebook in another, and a GPU box whose drivers nobody wanted to touch. That is precisely the failure mode FGCV was built to remove, and it is why she agreed to try it under a hard deadline.

Days 1–2: Audit the data before touching a model

The first decision point came fast, and it was not about architecture. Dana's team ran the existing labeled set through an ingestion pass to check class balance and duplicate frames. The audit surfaced two problems worth naming, because they are typical:

  • About 8% of "damaged" labels were actually glare on the shrink wrap — a labeling error rate that would have quietly capped accuracy no matter how long they trained.
  • Class distribution was skewed 11:1 toward undamaged cartons, which explained why the old model scored 96% overall and still missed most real damage.

Fixing those two things took a day and a half and cost nothing but attention. It is the least glamorous phase of any computer vision platform engagement and the one that most reliably determines the outcome.

Days 3–5: Train, evaluate, and resist the urge to tune forever

With a cleaned set, the team moved to training. The obstacle here was cultural, not technical: Dana's best engineer wanted to run a wide sweep of backbone architectures. The counter-argument was that the pilot's purpose was to prove the pipeline, not to win a benchmark. They settled on a small sweep — three configurations, evaluated on a held-out split with per-class recall as the headline metric rather than overall accuracy. The winning configuration hit 0.91 recall on damage at a false-positive rate the dock supervisor accepted after one calibration session.

Evaluation mattered more than training. Because runs were versioned against the exact dataset snapshot, Dana could show her operations director why a later experiment looked worse — it had been trained on a snapshot taken before the glare relabeling. That audit trail is the difference between a demo and a system, and it is the part teams usually build last, if ever.

Days 6–7: The deployment question that kills most pilots

The conveyor site had no cloud egress worth speaking of and a latency budget of roughly 120 milliseconds per frame. So the model had to run locally on modest hardware. This is where a lot of MLOps for CV tooling quietly gives up: quantization works in a notebook, then breaks in the container. Dana's team exported the model, ran it on a single mid-range GPU at the dock, and measured 38 milliseconds of inference per frame — comfortable headroom. A second camera stream pushed it to 71 milliseconds, still inside budget. They left the third stream on the old rule-based check as a fallback for the first week.

If you want the longer version of how that export-and-monitor path is structured, the team documented their rollout against the pipeline described in the platform's training-to-deployment workflow, which is worth reading before you commit to a hardware profile.

Days 8–9: Monitoring, and the first thing it caught

Deployment without monitoring is just a hope with a cron job. The team wired up alerts on input drift and prediction distribution, then went live. Within 36 hours, the system flagged a shift in average frame brightness on camera 4 — a lens that had collected dust. Recall on that stream had dropped about 6 points before anyone noticed visually. A wipe with a microfiber cloth restored it. That single catch justified the pilot to the finance team more than any accuracy chart.

Dana's summary of the whole run was blunt: "We spent two days on data, three on training, two on deployment, and two on the parts nobody budgets for. FGCV reports the pipeline is designed to take a model from labeled dataset to monitored deployment, and in our case that claim held up under a nine-day clock." We have no reason to doubt her; the numbers line up with what we saw in the run logs.

What we'd tell the next team

Three takeaways, in order of importance. First, audit labels before you audit architectures — the glare problem was worth more than any model choice. Second, pin evaluation to dataset versions or you will relearn the same lesson every quarter. Third, treat monitoring as a launch requirement, not a phase two. The pilot cost roughly nine engineering days and one mid-range GPU; the measurable result was a 0.91 per-class recall model running at 38 milliseconds per frame with drift alerts that caught a real failure in under two days. That is a modest, unglamorous win — which is exactly what production computer vision looks like.

See every redirect as a revenue event.

Book a live walkthrough of the MoeLink attribution dashboard. We'll route a real link through your stack in the call.

Get a live demo