Transcript
Sam: Welcome back to the SmartFactory Podcast Series. I’m your host, Sam Duchscherer. And once again, I’m joined by Ravi Jaikumar, Global Product Manager for SmartFactory Scheduling Solutions. Ravi, welcome back. First time I have had a repeat guest appearance.
Ravi: Thanks, Sam. Good to be back.
Sam: All right, so Ravi, last episode we talked about measuring the wrong thing. Today, I want to talk about something harder. I like to give you a little bit of a challenge here. Let’s talk about when you finally pick what looks like the right KPI. How do you know your scheduler is actually the reason it moved? And maybe we can bring in the concept from our last podcast; let’s use that X Factor example you mentioned. So, the scenario I’m thinking about is, suppose X Factor improves by 10%. Leadership is happy. Maybe customers are happy, too. But how do I know scheduling deserves the credit?
Ravi: I hate this question, but I also like this question. Most of the time, you don’t—not with the rigor you’d want before you spend money on the strength of it. And this is not going to be a popular answer for someone in my job to say out loud because I sell scheduling software, but it’s true. And I think the industry is better off saying it. Now, I do want to give you a real answer to that question, and I will. But I would like to do something first, because there is a step that people skip. Before you ask whether the scheduler moved the number, you have to be sure the number was worth moving. And in my experience, when someone shows me a 10% improvement, more often than not, the interesting problem isn’t the attribution. It’s the metric itself. So, do you mind if I come back to your question? I promise I’ll answer the question. Let’s sort out the KPI first, and then the attribution question gets much more answerable. Is that fair?
Sam: Oh man, you are leaving us with a teaser here! I love that. That’s why you’re my first guest that has been on my podcast twice. That is completely fair. Okay, so let’s just focus on the KPI side for a second, as you suggest. There’s one metric I hear all the time; let’s just maybe bring up moves—every fab seems to have it on a dashboard somewhere. Let’s talk a little bit more about that. Help me explain why.
Ravi: Because it’s the only thing that works at the time scale a shift actually runs on. Think about what a shift supervisor needs. They need a number at the end of the day or the end of the shift that tells them tonight went well, not next month, but tonight. Moves does that. It’s countable. It’s unambiguous. It updates in real time, and nobody argues about the definition. An operation either completed it or it did not. Other metrics like cycle time can’t do any of that. You don’t know a lot cycle time until the lot is finished. You can predict it, but you don’t know it, which could be weeks. You cannot run a shift on a number that arrives a month late.
So I want to be fair to moves. As a shift-level activity signal, it’s good and it’s legitimate. It’s doing a real job. The failure isn’t that fabs count moves. It’s that the shift metric quietly becomes the program metric—the number we improve against; the number in the business case for a scheduling software or a statement of work; the number in the quarterly review; that’s where it stops being that useful and starts being somewhat misleading.
Sam: Can you give me maybe another example so I can wrap my head around it a little bit better?
Ravi: Sure. Let’s imagine the last two hours of a shift at a bottleneck tool set. The operator has a choice.
Option A:
Option A is that there are six slots queued that all use the current setup. So run them back-to-back. No changeover, six operations completed before the shift ends. That’s option A.
Option B, change the setup, which costs you the better part of an hour, and run two lots that are behind schedule and feeding a downstream area that’s about to run dry and starve. So two operations completed.
So, what does the metric say? Option A looks better. It’s six better than two. But in reality, the fab actually needs option B, even though option A looks good in terms of moves. Because in option A, the two lots that mattered sat for another 12 hours and the downstream area starved. And now you’ve got an idle bottleneck tomorrow, which going back to the last episode is exactly the variability we said was a real enemy.
And here’s the part I want the listeners to hear: The operator didn’t do anything wrong. They optimized precisely the thing we asked them to optimize. They read the scoreboard and they played to it. So I stop blaming the floor and go back to the scoreboard. The floor is almost always rational. The scoreboard, or the benchmark, or the KPIs we give them isn’t.
Sam: So just to make sure that I understand, because last episode, I’m just thinking about how we discussed utilization is lying to us. So, answer this, is moves lying to us too?
Ravi: It’s the same lie in a different unit. Utilization asks how busy the tool was. The moves asks how much activity happened. Neither one has any concept of whether the right thing happened. They are both activity metrics wearing the costume of a performance metric. And they’re both popular for exactly the same reason. They’re easier to count, they’re easily available immediately, and they feel like productivity. So yes, it’s the same failure, just one level up. Last episode, it was the fab optimizing the visible number. This episode, it’s the improvement program doing it, which is worse because a program sets direction for years.
Sam: So, I just brought up moves as an example to kick us off, mainly because my first question, you had me pivot a little bit. But obviously moves isn’t enough. So why don’t we talk about other KPIs? What other KPIs should people be looking at?
Ravi: I would argue for three numbers. The first is an output metric, how the fab actually performed for the customer. That’s either on-time delivery or X Factor. On-time delivery is what the customer feels. X Factor is normalized to the work, so it doesn’t reward you for happening to run an easy mix. Ideally, you would want both; you want to watch both, but you at least need one.
The second is a constraint metric, and the obvious one is the Q-time compliance. That’s different in kind from the first. It isn’t something you optimize; it’s something you must not violate. When you blow up a queue time link, that isn’t a slow lot, that’s a scrap. So, it needs its own line not to be averaged into the performance score. So that’s the second one.
The third one is a trust metric, which is the one almost nobody has. How stable is the schedule and how often do people override it? Because a scheduler that is theoretically 3% better and reorders the queue every 15 minutes is worse in practice than the one that’s boring and predictable. People tend to plan around schedule, so if it crashes, they stop believing it. And then it doesn’t matter how good the math was.
So I’ll kind of like summarize it, right? Output, constraint, trust. If you have those three, then you can have an honest conversation. One number can always be gamed, and 12 numbers means that nobody’s accountable for any of them. But I think having a metric of three sets actually works.
Sam: So, you might have answered this already but let me ask in a different way to see if I get a different answer. So, if you were to walk into 10 fabs tomorrow, what’s the KPI you’d be most surprised that isn’t being tracked?
Ravi: It’s a different answer actually, schedule stability. Here’s the question I’d ask: Between the moment you published the schedule and the moment it was executed, how many times did it change and how far did it move? And not many people can answer that (or many fabs can answer that) and there’s no report for it. And yet, in every fab I visit, somebody tells me the floor doesn’t follow the schedule, which is a stability complaint dressed up as a compliance complaint.
Its close cousin is override rate. How often does a human look at the recommendation and do something else? Everyone has that data sitting in a log somewhere. Almost nobody reports on it. And I would go further. I don’t think override rate is a performance metric at all. I think it’s a prerequisite. If your operators are overriding 40% of the recommendations, then every other number you’re reporting is describing a system that’s actually not running. Does that make sense?
Sam: It does. It does make sense. And honestly, you sound like a proud parent of all these KPIs. But let me ask you a question here is, if I had to force you to pick one KPI, which one are you going to go with?
Ravi: I think I mentioned this in the last one, but I would like to reiterate because it’s important, it’s X Factor, because it’s normalized to the work. So, it’s the hardest one to accidentally flatter yourself with. But I want to push back on the question a little, because picking one is the exact failure we’ve been describing for the last two episodes.
Moves did not become a problem because moves is a bad number. Utilization is not a problem because utilization is a bad metric. It became a problem because it’s the only number. Anytime you make the single metric the most important and want to get it optimized in ways you didn’t intend, that’s not a fab thing. It’s a people thing. It happens in every organization. So, my real answer is X factor, on the condition that I also get to see the product mix next to it. Because without the mix, the X factor is just as gameable as anything else. And that’s a good bridge to back to your original question you started in the beginning, right?
Sam: Right! So we have the list of KPIs. Let’s go back to where we started. Suppose you have the ones you listed, and they all improve. How do I know scheduling specifically caused that improvement?
Ravi: Right. So this is where I owe you a real answer, starting with why it’s hard. You cannot AB test a FAB. There’s one FAB, one set of WIP. Every area is coupled to every other area. You can’t run the same quarter twice with a different scheduler and compare. And in that same quarter, a dozen other things moved. Tools went down, availability changed, product mix changed, the release rate changed, somebody qualified a second chamber, or the demand softened. So, the last one is a big one, by the way. If the loading drops, the X Factor improves on its own. So, a demand downturn makes any scheduler look brilliant. So what can you actually do? Four things, roughly in order of how much I trust them:
The first one is the phased rollout, which is the most practical (and much of the industry already does this), which is: deploy to one area or one tool set first, hold the others. You get a comparison in the same quarter under the same conditions. It’s imperfect because the areas are coupled and the untreated ones get contaminated. But it’s real evidence.
Shadow mode is my favorite for a different reason. You run the scheduler live, log what it would have recommended, but don’t act on it. Then you compare its decisions against what actually happened. That won’t prove ROI because the outcome you care about never occurred. But it’s the single best way to build trust before deployment. It costs you nothing operationally.
Then the third one, simulation counterfactual. Replay the period under the old policy in a validated model. This one’s legitimate, but be aware of what happens next. You stop arguing about the scheduler and start arguing about the model, simulation model.
And then finally, then the mixed normalization, which isn’t really a method, it’s hygiene. It applies to all of the above. An unnormalized before and after comparison is, I would say, the single most common way scheduling ROI gets misstated in the industry. And I mean in both directions. I’ve seen good deployments look flat because the mix got harder.
Sam: I’ve gotta be honest with you, though. This almost sounds like you’re telling listeners that ROI is impossible to prove.
Ravi: I can see why you might get that impression, but let me be precise here, because there’s a real difference. It isn’t unprovable. It’s attributable with stated assumptions. So those are different things in my mind. What you don’t get is a laboratory experiment, right? What you do get is a converging evidence—a phased rollout, plus a normalized comparison, plus a simulation view. And when they all point in the same direction, I think that’s a reasonable basis for a decision. That’s how most consequential decisions get made outside a lab.
What I’d actually be skeptical of is the opposite. If A vendor hands you one clean number with no assumptions attached, they haven’t actually solved the problem. They have made choices they’re not showing you. And here’s the practical suggestion I’d give anyone who’s starting a deployment. Agree how you’re going to measure before you turn it on. Pick the baseline window, pick the normalization, pick which areas are treated and which are not, and write it down while nobody knows the answers yet. Because if you wait until the results are in and then start negotiating the methodology, you will never agree. Everyone will have a number that supports the position they already had. I watched that happen before and it’s a bad quarter for everyone.
Sam: Fair enough. I hear you. I love to learn from your experiences. So, thank you for sharing that. Let’s wrap it up, though, with some myth versus reality statements. I feel like you had really good answers last time, so I wanted to give you some new statements to enlighten us. And like last time, they’re not too hard. Maybe you’ve already answered them, but they’re good summary points. So, myth versus reality. Let’s go through it. Let me get my first one here in my notes. All right, so the first one: you should measure a scheduler against the KPI the FAB already tracks.
Ravi: You do have to speak the fab’s existing language. If you walk in and tell a team that everything they measure is wrong, you’re not going to get a second meeting, and you probably don’t deserve one. And I would say this claim is half true. So that’s the true half, which I just mentioned. The wrong half, the incumbent metric is always usually a big part of why the problem exists. If moves is what the fab tracks, then measuring your scheduler against moves guarantees you optimize the exact thing that’s hurting them. You’ll win the evaluation, but you’ll lose the fab. So, this is what I would actually do, is add the missing output metric alongside the incumbent one. Don’t replace anything, run both for a quarter and let people watch them diverge. That divergence is a far better argument than anything I could say in a slide because they found it in their own data.
Sam: Oh, okay. All right, this one you might have already answered, but I’m going to ask it anyway. And it goes back to this X Factor, your favorite one. If X Factor improves, the scheduler is working. Is this a myth versus reality statement?
Ravi: This might sound that I’m contradicting myself because I praised X Factor as a metric so much in the first two episodes, but it’s a myth. And this is the one I would most like people to take away. X Factor improves whenever the fab gets less busy. So go back to episode one. Queue time is driven up by loading. So drop the loading and the queue falls on its own with nobody doing anything clever. So, a soft quarter makes every scheduler in the world look like a genius. And the reverse is worse. During a hard ramp, a scheduler can be doing genuinely excellent work. One, the X Factor degrades because loading went up faster than the improvement, which means the number can move in the right direction for the wrong reason and in the wrong direction for the right reason. That’s why I said: I want the mix and the loading sitting next to it. X Factor on its own is a weather report. X Factor with loading and mix is like a diagnosis.
Sam: Oh, okay. I’m glad I asked that one, then. I thought it was a little bit redundant, but your answer gave me some more things to think about. Love that. All right, the last one on myth versus reality. And this one might be easy as well, but we’ll see where we land. It is the KPI with the nicest trend line is usually the most useful.
Ravi: This one’s very easy, Sam, and honestly, it’s closer to the inverse. A beautifully smooth trend line usually means one of two things: Either the metric is aggregated across so many products, areas, weeks, all that’s interesting variation that’s been averaged out of it. Or it’s a metric that simply isn’t very sensitive to what you’re doing. Fabs are not smooth. Tools go down, mixes shift, hot lots arrive. A metric that faithfully reflects that reality is going to be a bit noisy. So that’s a feature of the fab. The nice trend line is popular because it presents well, it goes into the presentation deck. Nobody asks hard questions about it. So, my rule of thumb is simple. If a metric never surprises you, it isn’t telling you anything. You’re not measuring, you’re decorating.
Sam: Oh man, see, I had to end it on an easy one so that you would come back for maybe a third episode. But Ravi, thanks again for joining this time.
Ravi: Thank you, Sam. I appreciate it.
Sam: And thanks to everyone listening. My summary is: The easiest number to measure isn’t always the one that creates value.
And with that, until next time, keep learning, keep questioning assumptions, and keep looking beyond the metric everyone else is celebrating.
