Skip to content
TrackPodcasts
newsSep 15, 20261:41:43pending

A Machine Goes Down. How Should Your Production Plan React?

About this episode

A machine goes down. Maintenance receives the alarm. The operator stops production. The MES records an interrupted operation. But your ERP may still believe that the machine has capacity. And your production schedule may still promise that every order will be completed on time. That is where a machine breakdown stops being only a maintenance problem and becomes a production planning problem. In this episode, we explore what should actually happen to a production plan when a critical machine becomes unavailable — from detecting the disruption and estimating lost capacity to identifying affected orders, evaluating alternative resources, rescheduling with finite capacity, and releasing a production plan the factory can actually execute.

NOT EVERY MACHINE STOP REQUIRES REPLANNING
Modern factories generate enormous numbers of operational signals. PLCs report faults. IoT systems detect machine-state changes. MES platforms track interrupted operations. Maintenance systems record incidents. But a machine stopping for a few minutes does not automatically mean the entire production schedule should change. A short interruption might be caused by an operator clearing a jam. It might be a normal tool change. The machine could simply be waiting for material. Or a sensor could report a state that does not represent the actual production situation. The important distinction is between a machine signal and a planning event. A production plan should react when the interruption creates a meaningful loss of capacity that normal shop-floor recovery can no longer absorb. That threshold depends on the production environment. A ten-minute stop on a resource surrounded by large buffers may have almost no impact. The same ten-minute interruption on a bottleneck resource producing a make-to-order component with a tight customer deadline could require immediate attention. Good production planning therefore does not react to every signal. It reacts to operational consequences.

HOW LONG WILL THE MACHINE REALLY BE DOWN?
Once downtime becomes significant enough to affect production, one question becomes critical: How long will the capacity be unavailable? Unfortunately, maintenance rarely knows the exact answer immediately. A technician may know that a drive has failed but not whether a reset will solve the problem, whether a component needs replacement, or whether additional safety checks will be required. Production planning should therefore avoid building an entire schedule around an uncertain repair timestamp. Instead, planners can work with a recovery window and confidence level. For example: A short outage may require no schedule change. A medium outage may require selected orders to move. A full-shift outage may require active rescheduling. A longer outage may threaten customer commitments and require escalation. This turns maintenance information into a useful planning input without pretending that an early repair estimate is a guarantee.

A MACHINE IS MORE THAN AVAILABLE HOURS
One of the biggest mistakes in production planning is assuming that another machine with an empty calendar automatically provides replacement capacity. It does not. A resource needs the correct capability. The right fixture may be required. A qualified production program may need to exist. An operator with the necessary skills must be available. Quality may need to approve the alternate process. Material needs to be physically available. Tooling and inspection capacity may also be required. This creates an important distinction: Machine availability is not the same as executable production capacity. An empty four-hour window on another machine means very little if the product cannot actually be produced there. Production planning therefore needs a model connecting Product, Process and Resource. The product defines what needs to be manufactured. The process defines the operations required. The resource defines where those operations can actually run and under which constraints. That context becomes essential when production is disrupted.

WHICH ORDERS ARE ACTUALLY AFFECTED?
Once a machine failure is confirmed, planners need to identify the work that genuinely depends on that resource. Start with the order currently running. How much has already been completed? What quantity remains? Can the partially processed material safely wait? Does interrupted work require inspection, rework, or scrap approval? Then examine the queue. Which orders are physically waiting? Which orders are released? Which ones have tight downstream dependencies? Which have alternative routings? Which rely exclusively on the failed resource? Not every order scheduled on the machine carries the same risk. Some may easily move. Others may have enough delivery buffer to wait. Some may depend on a customer shipment. Others may feed a critical assembly operation. And some may have no realistic alternative at all. The goal is not to declare every order an emergency. The goal is to identify real production exposure.

THE BOTTLENECK CAN MOVE
Suppose Machine 4 fails. Machine 6 has available capacity. Moving urgent orders to Machine 6 appears to solve the problem. But what happens next? Perhaps Machine 6 feeds an inspection station that is already operating near capacity. The orders move successfully — and simply create another queue somewhere else. A machine breakdown can therefore move the production constraint. The new bottleneck could become another machine, an inspection station, a qualified operator, a furnace, a fixture, a transport resource, or even a tool. This is why production planners need to evaluate the entire local production flow, not simply find another empty machine slot. Moving work can move the problem. A good rescheduling decision considers what happens downstream and upstream after every significant schedule change.

WHAT SHOULD THE PRODUCTION PLAN PROTECT?
When capacity disappears, not every order can necessarily remain exactly where it was. Production therefore needs clear priorities. Customer delivery commitments matter. But so do internal milestones. Safety stock can matter. Material shelf life can matter. Campaign rules can matter. Setup costs can matter. Quality requirements can matter. An urgent-looking order may require a long changeover and material that has not yet arrived. Another order with a slightly later due date may already have material staged and require almost no setup. Blindly sorting everything by due date can therefore produce a worse production schedule. Priority rules should exist before the machine breaks down. Otherwise, disruption planning becomes a competition between whoever calls first, whoever complains loudest, and whoever has the most senior manager copied into an email. Production planning needs policy, not panic. 

POSSIBLE CAPACITY VS EXECUTABLE CAPACITY
Before moving an order to another resource, planners need to test the full chain of constraints. Can the machine technically perform the operation? Is the tooling available? Is a qualified operator available? Is the material physically ready? Has quality approved the alternative? Does the production sequence allow the change? Are batch rules respected? Does shelf life create another time constraint? Has maintenance actually released the resource for production? A resource can be theoretically capable without being operationally available. This distinction between a possible move and an executable move is fundamental to realistic manufacturing scheduling. A possible move says that another machine could perform the operation. An executable move means that machine, tooling, operator, material, process approval, sequence, safety requirements, and available time all align. Only then does that empty slot become real capacity.

BUILD SCENARIOS INSTEAD OF SEARCHING FOR ONE PERFECT ANSWER
Machine downtime introduces uncertainty. Trying to calculate one perfect replacement schedule can therefore create false confidence. A better approach is to generate a small number of realistic response scenarios. One scenario might keep the work on the failed resource and wait for repair. Another might move selected orders to approved alternative resources. A third could use overtime or an additional shift. Another possibility might involve splitting an order where production and quality rules permit it. Each scenario should explain: What changes? Which orders move? Which customer commitments are protected? Which setups are added? Which resources are required? What assumptions does the scenario depend on? And what happens if the repair takes longer? The objective is not to create dozens of simulations. Usually, production needs a preferred response, a fallback, and perhaps a more aggressive option if delivery risk becomes unacceptable. 

WHY EXCEL REPLANNING BREAKS UNDER PRESSURE
Excel remains extremely useful for local analysis. But problems begin when the spreadsheet becomes the new production schedule while the actual production environment continues changing elsewhere. Maintenance updates the repair estimate. A supervisor moves an order. Customer service changes a priority. MES records new production progress. Material moves. Another resource becomes unavailable. Suddenly several people have several different versions of the production plan. And everyone believes their version is correct. Then come the familiar filenames: Final. Final_v2. Final_v2_revised. The factory version of archaeology. The underlying problem is not Excel itself. The problem is the absence of a shared, controlled source of truth. Production needs one released schedule that clearly shows which facts were used, which orders changed, when the schedule was calculated, and which version the shop floor should actually execute.

Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.

Get every episode summarized

Each time M365.FM - Modern work, security, and productivity with Microsoft 365 publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Hosts & guests

No transcript yet

This episode has not been transcribed. Request it and it moves to the front of the queue.

A Machine Goes Down. How Should Your Production Plan React?

M365.FM - Modern work, security, and productivity with Microsoft 365

0:00
1:41:43

More episodes

More from M365.FM - Modern work, security, and productivity with Microsoft 365

View all episodes →