A 24-Robot, Two-Warehouse Rollout Needs One Commissioning Scorecard

Buying mobile robots is a technology decision. Proving they work across different warehouses is an operating decision.
O'Neill Logistics provides a timely example. Modern Materials Handling reports that the 3PL plans to deploy 24 collaborative mobile robots across facilities in Monroe, New Jersey, and Savannah, Georgia, with go-live planned for the fourth quarter of 2026. The New Jersey operation supports retail and direct-to-consumer fulfillment, while the one-million-square-foot Savannah campus handles omnichannel orders.
Those are not interchangeable environments. Different layouts, order profiles, labor mixes, and cutoff times can make the same fleet look excellent at one site and disappointing at another. A single commissioning scorecard creates a common standard without pretending the buildings are identical.
Separate acceptance from operational value
Vendor acceptance testing should answer whether the system supplied is functional. Can each robot navigate its approved zones, avoid obstacles, receive work, complete a mission, charge correctly, and exchange data with warehouse systems? Do emergency stops, alerts, user permissions, and recovery procedures behave as designed?
These are pass-or-fail conditions. They belong at the top of the scorecard because an operation should not trade safety or basic functionality for a higher pick rate.
Operational measures answer a different question: does the installed system improve the warehouse? O'Neill expects the robots to automate repetitive material movement, reduce unproductive walking, and support picking, putting, transport, and mobile sorting. Commissioning therefore must measure those outcomes rather than treating a successful demonstration route as proof of value.
The distinction matters because robot performance can appear healthy while the process around it fails. A unit may complete 99% of assigned missions, yet associates may wait for work, aisles may become congested, or packing may be overwhelmed by faster picking. Acceptance proves capability; operating metrics prove usefulness.
Build the scorecard in five layers
1. Safety and control
Record collisions, near misses, emergency stops, blocked egress incidents, unsafe manual interventions, and time required to acknowledge and resolve alarms. Include observations by shift and zone, not only system-generated events. Zero recordable incidents is essential, but near-miss reporting reveals risks before they become injuries.
Training completion and recovery competence belong here too. Measure the share of designated associates and supervisors who can pause, redirect, restart, and safely isolate a robot without vendor assistance.
2. Flow and travel
Establish a manual baseline before go-live: associate walking distance per order or line, cart travel, queue time, and dwell at pick and put locations. After deployment, compare equivalent order cohorts. The relevant result is not robot miles traveled; it is avoidable human travel removed without moving delay elsewhere.
Track congestion as minutes of blocked or slow travel per robot-hour, segmented by aisle and time window. A building-wide average can hide a severe pinch point during carrier cutoff. Heat maps and exception logs should identify whether the cause is layout, replenishment, task sequencing, or human-robot interaction.
3. Productivity and service
Measure lines and orders per direct labor hour, picks per hour, order-cycle time, cutoff attainment, accuracy, and rework. Use medians and percentiles as well as averages so a strong daytime shift does not conceal unstable peak performance.
External benchmarks show why productivity deserves attention but should not become a promise. Inbound Logistics documented an AMR implementation that increased picking from 30 to 110 picks per hour, a reported 266% increase after a 90-day implementation. MHI separately cites a company that achieved a 200% productivity increase after adding AMRs. These examples establish potential, not a universal target. O'Neill's scorecard should compare each site against its own controlled baseline.
4. Reliability and support
Track fleet availability, mission completion, mean time between failures, mean time to recover, charging time, interventions per 100 missions, wireless disruptions, and integration errors. Define uptime carefully: a powered-on robot waiting because it cannot receive work is not operationally available.
Every failure needs an owner and a clock. Separate robot hardware, software, facility network, WMS interface, master data, and process causes. That classification prevents every stoppage from being labeled a generic “robot issue” and shows where corrective investment belongs.
5. Workforce and adoption
Measure training time, associate feedback, fatigue indicators, voluntary workarounds, and tasks returned to manual handling. MHI notes that AMRs can reduce repetitive walking and physical strain; its cited case also connected automation with improved retention. Those gains are real only if workers use the designed process and trust its safety.
Normalize without erasing site differences
The two facilities should use identical definitions, observation windows, severity levels, and calculation rules. But targets should reflect each site's baseline and operating profile.
Report improvement percentages beside absolute values. Compare direct-to-consumer orders with similar orders, not with store-replenishment cases. Segment by order lines, cube, zone, shift, and demand band. For Savannah's larger footprint, travel reduction may carry more weight; for a dense operation, congestion and handoff time may dominate.
A useful scorecard displays three views: the raw site result, change from that site's baseline, and variance from the approved target. This makes results comparable while preserving the context needed to act on them.
Turn Q4 exceptions into rollout gates
During go-live, log every exception with timestamp, site, zone, shift, order type, robot, symptom, operational impact, root-cause category, workaround, owner, and closure evidence. Review critical safety and service exceptions daily; review trends weekly.
Broader deployment should require explicit gates: no unresolved critical safety defects, stable integration performance, acceptable peak-period congestion, target service levels, trained local support, and sustained productivity improvement over a representative demand cycle. One strong demo shift is not enough.
The best commissioning decision may be “scale,” “scale after correction,” or “hold.” A shared scorecard makes each outcome defensible. It turns 24 robots in two warehouses from parallel installations into one controlled learning program.
Ready to connect warehouse execution with transportation planning and exception visibility? Request a CXTMS demo to see how one operational platform can keep fulfillment and freight decisions aligned.


