Episode 6

Why Your Bottleneck at 95% Is Lying to You with Ravi Jaikumar

Transcript

Sam: Welcome to the Smart Factory Podcast Series. I’m your host, Sam Duchscherer, and today I’m joined by Ravi Jaikumar, the Global Product Manager of Smart Factory Scheduling Solutions. Ravi, welcome.

Ravi: Thank you, Sam. Glad to be here and thank you for doing this, though I’ll warn you up front, you picked a topic I can talk about for a long time.

Sam: Yeah, yeah, before this, we were discussing offline how maybe for you, this episode actually might turn into one of many. And I’m really excited about that.

Ravi: Yeah, once you start pulling on scheduling and AI, that’s a lot of thread there. But if I had to pick the one idea that runs underneath all of it, it’s this: In manufacturing, we optimize what we can easily measure. And I want to be careful here because that’s not a criticism. It’s completely rational. You manage what shows up on the report, and you report what you can count reliably at the end of a shift. The problem is that the most visible metric is not usually the one that actually matters. It’s just the one that’s easiest to count. So, you end up with a fab full of genuinely smart people all working hard, all hitting their numbers. And the thing the customer actually cares about is quietly getting worse. That’s the story I’d like to tell across a few of these. Utilization is just the cleanest place to start.

Sam: So, let’s break down your vision. And today, let’s just focus on the topic that sounds straightforward, but does create a lot of debate. And you already mentioned it—utilization. Most managers want more of it, and most dashboards put it front and center. But in your experience, does higher utilization actually mean better fab performance?

Ravi: Good question. Up to a point, yes. Past that point, no. And the point comes earlier than most people expect, actually. Here’s the core of it. Utilization answers one question: How busy is my tool? It doesn’t answer the question the customer is asking and the customer is paying for, which is how long did my lot wait? Those feel like the same question, but they are not. And at high utilization, they don’t just differ; they actively pull apart.

Sam: So, let’s dive a little bit into that. Could you put it into an analogy that I could use to explain to someone who knows nothing about semiconductor manufacturing?

Ravi: Okay, think about a highway—a road at 50% capacity and a road at 95% capacity. They’re both moving cars. From a helicopter, they look like they’re both working. And honestly, the busier one looks better, fuller, more productive. Now, one truck taps its brakes. At 50%, the car behind changes lanes 30 seconds later. It’s like nothing happened. The road absorbs it. But at 95%, there is no way to go. That one tap becomes a wave that travels backwards for miles. An hour later, there are people sitting completely still who have no idea that the truck was ever involved. Same disruption, completely different outcome, and the only thing that changed is how full the road was. A fab is that road. A truck tapping its brake is like a tool going down, or a hot lot cutting in, or a batch that didn’t form on time. And at 95% utilization, your fab has no lanes to change into.

Sam: I could pick at you a little bit. An idle tool is wasted capital. And, you know, that’s just in my opinion, but isn’t the whole purpose to keep it busy?

Ravi: You’re actually right, and that’s exactly why this is such a hard habit to break. These tools cost tens of millions of dollars. Nobody wants to walk the floor and see one sitting idle. It feels like burning money. And there’s a real argument there; I’m not going to pretend otherwise. But here’s the reframe I’d like to offer: You’re not choosing between a busy tool and an idle tool. You’re choosing where the waiting happens, because the variability is there either way. Tools go down; lots arrive in bursts; you don’t get to opt out of that. So, either the tool absorbs it—that’s a little idle time—or the WIP of the factory absorbs it and that’s queue time. Same disruption; you’re deciding who holds it. And here’s what makes it asymmetric: idle tool time is a known bounded cost. You can put a number on it this afternoon. Queue time compounds and it shows up months later as inventory, late deliveries, slower yield learning, usually in someone else’s budget. So, we pay very close attention to the cost we can see and almost none to the cost we cannot.

Sam: Well, I think you have painted the simple version for me. Actually, I don’t think; I know that. That was really great. But give me the real version.  The story that fab managers now that they have this semiconductor knowledge can relate to.

Ravi: All right. There’s a piece of queuing theory called Kingman’s equation. Every industrial engineer and every fab actually knows it. They study that in school. They implement that in their factories. But I’ll say the name once and never again. What it says is that the Kingman’s equation on the queuing theory is that the time a lot spends waiting is essentially basically three things multiplied together. First one is variability, second one is it’s a utilization term, and the third one is process time itself.

The one that catches people out is a utilization term. It’s basically the formula is u over 1, minus u. U stands for utilization. So just run the numbers, right? For example, at 80% utilization, that term in the formula is 4, right? So it’s like 80 divided by 100, minus 80. So, the term is 4. At 90%, it’s 9. At 95%, it’s 19. At 98, it’s 49. So, look at the step from 90 to 95% utilization— how that term, the value of the term changes. Ninety to 95% utilization is just five percentage points, right? But the term goes from 9 to 19. We have roughly double the queue. That’s the whole disconnect here. The manager reads five points. The physics delivers double.

But if you ask me where I’d actually spend my energy, it wouldn’t be on that particular term. It will be the first one. That’s called variability. Because variability and utilization are multiplied together, which means they are substitutes. A very smooth fab can run at 95% and still be fine. A lumpy fab can’t survive at 85%, so that lumpiness and the smoothness is determined by the variability. And the lumpiness comes from two places; the variability comes from two places. One is arrivals—releases that come in batches, lot arrivals, a shift change that dumps WIP into the line, hot lots injected on top of a plan that didn’t expect them and didn’t plan for them, and the process itself, unscheduled downtime (which in most fabs is the single biggest contributor by a distance), setups, qualifications, waiting to form a batch, full batch, rework loops… So when a fab tells me they want more out of the bottleneck, my first question isn’t about dispatch rules. It’s how variable is that tool set right now? Because at 95% utilization, having your variability half, like reducing it by 50%, will buy you more cycle time than any dispatch rule I could hand you or write for you.

Sam: So, if I’m understanding you, is the answer just to run the bottleneck at around 85% and accept it? I just feel like no CFO is going to sign up for that?

Ravi: No, you’re right. Nobody’s signing up for that. I’m not saying run the bottleneck cooler. I’m saying you have to earn the right to run it hot. Go back to the equation, variability times, the utilization term, right? So, if you cut variability in half, you can hold the same cycle time at a meaningfully higher utilization. The higher number becomes safe. So, the mistake was never running hot. The mistake is running hot while staying lumpy and then being surprised by the result. And the reason it keeps happening is organizational. In most fabs, utilization and variability are two separate programs owned by two separate teams and reviewed in two different meetings. They are not separate. They are not different. They are two terms in the same product.

Sam: Okay, okay, that makes more sense. Then make the case to a CFO. Why should anyone trade a utilization point for cycle time?

Ravi: Honestly, the answer changes depending on who’s in the room, Sam. And let me do a few. Start with the CFO or the person writing the check for buying new tools, because the language there is inventory. There is a simple rule called Little’s Law. Again, every industry engineer in the fab knows it. WIP equals throughput time cycle time. Read it backwards. And it says cycle time is inventory. Cut cycle time 20% at the same output and you’re carrying 20% less WIP. Advanced notes, that is an enormous amount of capital sitting on the floor doing nothing. That’s not an efficiency argument. That’s a balance sheet argument.

Now, the process time and yield people. For them, cycle time is a learning rate. Every trip through the line is one feedback cycle from electrical test back to process. Shorter cycle time means more learning per quarter. During a ramp, that is the entire game; the fab that learns fastest gets to yield first and getting to yield first is worth more than any utilization point you will ever recover.

For whoever owns the customer relationship, it’s about commitment. Long cycle times mean you’re locking in a product mix months before you know what demand actually looks like. Short cycle time is optionality. You get to decide closer to the truth. And then there’s the one that never makes it to the business case—queue time constraints. Some links in the flow have a hard clock on them. We call them time-sensitive operations. Longer queues mean more loads sitting near that limit. And when you blow it, that isn’t a delay. That’s a scrap. That’s a rework. You’ve paid for every step up to that point, and you get nothing back if the lot goes to scrap.

So, the framing I leave a CFO is with this: Utilization is an input; cycle time and on-time delivery are the outputs. We manage the input because it’s the one we can see, and we are paying for that convenience somewhere else on the P and L. Hope that makes sense.

Sam: This is why I love to have people like yourself on this podcast. That was a great explanation. I personally learned something and I would hope that our audience feels the same way. It was such a great explanation when you went through the different personas there.

So, let’s transition a bit though. I typically, if you follow my series, like to wrap up these podcasts with a lightning round. But for you, I want to change the rules a bit. Instead, we discussed doing a myth versus reality segment. So, for the ending of this podcast, I like to throw out some things people commonly believe, and you can help separate facts from fiction. Are you ready for that challenge?

Ravi: Let’s do it, Sam. Go ahead.

Sam: All right, so the first one is if your bottleneck tool set reaches 95% utilization, have you achieved peak fab performance?

Ravi: Easy: myth. And it’s a comfortable one because the number itself is real. The tool genuinely is busy. Here’s how it actually plays out: A fab pushes the bottleneck from 90 to 95%. The number goes out, the best utilization in the plant’s history, and it’s earned. People worked hard for that.

Week one, nothing looks wrong. The queue in front of the tool set grows, but WIP is a slow-moving number, and nobody panics about a slightly longer line. Week two, the queue is deep enough that when a tool goes down for four hours, there is no time left to recover with. That four hours doesn’t get absorbed anymore. It just gets added.

Now, let’s look forward. Week number three: the hot lots start missing, so people start expediting, which means cutting in line, which means more variability, which makes the queue worse. That’s the part that really hurts. The response to the problem feeds the problem. And the entire time the utilization looks fantastic because the tool is busy. It’s just busy on a queue that’s 40 lots deep. So, the honest version is that at 95% isn’t peak performance. It’s actually peak fragility. We have removed all your margin for error, and a fab is a machine that produces errors.

Sam: All right, all right. That one was easy. It’s your first time on the podcast. I decided to go a little easier on you. So, let’s try a more challenging one. The next one is more automation always reduces cycle time.

Ravi: I would say that’s half myth. And again, I’m going to assume a definition of automation for this question, right? It’s half myth, and it’s the half that gets expensive. Automation, in my understanding, moves material faster. That’s true, and it’s measurable. But transport time is usually a small slice of the total cycle time. The vast majority of a lot’s life is spent waiting in a queue, not moving between the tools. So, automation only helps cycle time if it reduces variability. And there is good news; the good news is that it often does. It’s consistent. It doesn’t take breaks. It doesn’t get distracted at shift change—but it can go the other way. An AMHS with availability problems adds a brand-new source of downtime and a brand-new source of variability. A rigid delivery or batching policy can force a lot to wait for a scheduled move instead of going now. And here’s the trap: In both of these cases, your transport metrics improve—delivery times, moves per hour, all better. And the cycle time gets worse. You’ve automated the 15% and made the 85% harder. So, the accurate statement isn’t automation reduces cycle time. It’s reliable automation reduces cycle time. The word doing all the work in that sentence is reliable.

Sam: I couldn’t agree more. So, throughout this podcast, you’ve been hinting that utilization might be not the only thing to consider.  So, the final myth versus reality is, if utilization is the wrong number to celebrate, what’s the right one?

Ravi: Ah, now we get to the right question, and it’s a bigger one than we have time for here. But the short answer is, you want a number that’s normalized to the work, not to the tool. Utilization measures the tool. What you actually want to know is how long a lot took compared to how long it should have taken if it had never waited for anything. So, in our world, that’s called X factor. So basically, the answer to your question is cycle time. And there are different ways to measure cycle time, and there are different KPIs that measure cycle time all in a different way.  But this is why exactly it needs its own episode. Choosing the metric turns out to be the easy part. The hard part is proving your scheduler is what moved it, right? You can’t AB test a fab. You’ve got one fab, one set of WIP, and everything is coupled to everything else. I could go into more detail, but I’d rather do that properly next time than rush it now. But the answer to your question is cycle time, and it can be expressed in multiple ways.

Sam: Yeah, Ravi, it definitely sounds like we have many more discussions to come. But for now, I appreciate your time.

Ravi:Thank you, Sam. And this was fun. And genuinely, if one person listens to this and goes and looks at their variability before they chase another utilization point, that’s a good outcome for us.

Sam: Absolutely. Yeah. For our listeners, keep learning, keep questioning assumptions, and stay tuned for the next episode.

Ravi: Thank you.

About the Author

Picture of Samantha Duchscherer, Global Product Manager
Samantha Duchscherer, Global Product Manager
Samantha is the Global Product Manager overseeing SmartFactory AI™ activities. Prior to joining Applied Materials Automation Product Group Samantha was Manager of Industry 4.0 at Bosch, where she also was previously a Data Scientist. She also has experience as a Research Associate for the Geographic Information Science and Technology Group of Oak Ridge National Laboratory. She holds a M.S. in Mathematics from the University of Tennessee, Knoxville, and a B.S. in Mathematics from University of North Georgia, Dahlonega.