Skip to content
Five ChipLLC

Case study

Run-to-run control for low-volume products in a wafer fab

Most of the fab's products ran too rarely for automated process control to learn from them, so they were reworked until they passed. A Six Sigma project and a small change inside the control system I had built let them borrow the high-volume product's data. Twenty years on, it's still running.

Client
Seagate Technology, wafer fabrication
Role
Process control system architect, Six Sigma project lead
Timeline
Baseline February 2005. In production June 2005. Still running.
Stack
Java, Enterprise JavaBeans, Oracle, XML over HTTP, JUnit

Define

Most of the fab's products were small experimental R&D runs.

In a wafer fab, every layer has to line up with the one beneath it. The stepper aligns a photomask to marks on the wafer, and the offset it needs changes a little with every lot. Automated run-to-run control learns that offset from recent lots and recommends settings for the next one. It works beautifully when a product runs many lots a day.

About 80 percent of the products at this fab were experimental products that ran less than one lot a day on a given tool-stage. For them there was no recent history to learn from, so their overlay missed specification more often, and a lot that missed was reworked until it passed. Rework ate turn time on development lots, stepper and metrology capacity, and engineering hours spent working out corrections by hand. It had been going on for nine months, since a new process arrived on the steppers.

The year before, I had architected the fab's Advanced Process Control system and worked with engineering and IT to put it into production. This project was about making it serve the products it couldn't yet see.

Discovery

The photo engineers already knew where the settings came from.

The team was four photo engineers and me. We started with a mind map of everything that could move overlay at this stage, then a system map of suppliers, inputs, process, outputs and customers, with the feed-forward and feedback loops drawn in. The most useful finding came from sitting with the engineers who set up lots every day: the overlay settings for a low-volume lot could come from any of three places, and none of them was good enough.

  1. 01

    Table defaults

    Stored in the factory database. Correct on the day they were entered, and quietly out of date after that.

  2. 02

    Engineering

    A photo engineer working out the correction by hand. Right, but there weren't enough engineers to do it for every lot.

  3. 03

    The run-to-run controller

    Automatic, and excellent for the high-volume product. For a product that ran once a week, it had nothing recent to learn from.

Two key inputs fell out of that: the overlay settings used at this stage, and the alignment marks created at the earlier stripe stage. Everything after this point was about proving which one mattered, and by how much.

Measure

Prove the ruler before trusting the data.

Before baselining anything, a gage study on the overlay metrology tool showed that less than four percent of the process variation was measurement error. The numbers could be trusted.

Then twelve weeks of production data set the baseline. In the direction that mattered, the high-volume product was capable. The two low-volume products were not: their process capability index was well under one, which means a meaningful share of lots would land outside specification.

Baseline overlay-y capability, February to April 2005
ProductLots a dayOverlay-y CpkOverlay-y sigma
High-volume productSeveral1.15139 nm
Low-volume product AUnder one0.71232 nm
Low-volume product BUnder one0.57244 nm

Goal for the low-volume products: overlay-y sigma at or below 153 nm, within 10 percent of the high-volume product.

Analyze

Two causes, two owners.

Three candidate causes went under the microscope: gaps in feedback data, which of two steppers patterned this stage, and which of two tools patterned the earlier stripe stage. Tracing each high-volume lot's path through both tools and testing the variances showed that the stepper at this stage mattered, with roughly a four-nanometre difference between the two machines, and that the stripe tool barely registered.

The data gaps refused to correlate with anything, and the reason was instructive: engineers were manually overriding the settings so often that the effect of missing data was hidden underneath their corrections. That constant manual intervention was itself the strongest evidence that the gaps mattered. So the work split cleanly. Photo engineering took the stepper-to-stepper variance. I took the data.

Improve

Let low-volume products borrow the high-volume product's data.

The control system already did the right thing for the high-volume product: every lot's measured overlay became input for the next lot's recommendation. The low-volume products were failing only because that loop had nothing recent to work with. The fix was to let related products share a pool of recent lots, under rules the photo engineers set.

Run-to-run control loop: the tool controller asks the control system for settings, the control system recommends them from recent lots, the stepper runs the lot, the metrology tool measures it, and the data uploader feeds the measurement back for the next recommendation. One lot at a time, every lot, around the clock.Tool controllerA lot arrives at the stepper. Asks for settings.Control system: recommendFrom recent lots of this product, y = f(x)StepperRuns the lot with the recommended settingsMetrology toolMeasures overlay on the finished wafersData uploaderNew records go to the control system as XMLFeedbackMeasured overlay becomesinput for the next lot.With groupingThe next lot of any productin the group can use it.
The run-to-run loop as it already worked for the high-volume product. Grouping changes only what the recommendation step is allowed to look at.
Product grouping: the high-volume product shares its recent lots into a group but never uses the group's data; each low-volume product both shares into and uses the group, so its recommendations draw on many more recent lots. One group per tool-stage, configured by the photo engineers.High-volume productMany lots a dayShares. Never uses.Low-volume product AUnder a lot a dayShares and usesLow-volume product BUnder a lot a dayShares and usesGroup: recent lotsOne row per productand stage, two flags:share lots to groupuse group lotsA recommendation for alow-volume lot now drawson the whole group.
The grouping design. Two flags per product and stage keep the high-volume product isolated while the low-volume products draw on the whole group.

Requirements the engineers set

  • A low-volume product must be able to draw feedback data from a group of related high- and low-volume products.
  • The high-volume product must not be affected. It keeps using only its own lots, exactly as before.
  • Photo engineering must be able to configure and maintain groups themselves, without a developer.

How it was built

  1. 1

    A parameters database for grouping

    One row per product and stage, with two flags: share lots to the group, and use lots from the group. Edited in the control system's existing web interface.

  2. 2

    Controller logic changes

    The run-to-run controller reads the group, loops through its members to find who is sharing, and builds the query that retrieves the pooled lot history. It also checks for configuration errors such as a duplicate entry for the same product and stage.

  3. 3

    Tests before production

    Ten new tests covered grouping during measurement submissions and recommendations, groups with and without entries, and the no-lots-found case. To pass, every one of the 179 tests already in the suite had to pass too, unmodified.

With the tests green, the modified controller went into production for a month-long evaluation on live product in June 2005, one group of three products at one stage.

Control

Control designed in, not bolted on.

Because the change lived inside the control system, the control plan was mostly the system's own behaviour, extended carefully rather than replaced.

  • Alerts designed in

    Two new error codes for grouping configuration mistakes, wired to a distribution list owned by photo engineering.

  • Fail safe, not stop

    If a group is misconfigured, the controller falls back to a standard, ungrouped recommendation so no work in progress is put on hold.

  • Nothing else changed

    The existing SPC charts, out-of-control action plans and no-past-lots alerts were untouched and verified by tests, so the shift teams needed no retraining.

  • Training for the engineers

    Photo engineers were shown the grouping screen and could create and change groups on their own. Group ownership stayed with them.

Results

One product met the goal. One came close. The high-volume product didn't move.

Overlay-y before and after grouping
ProductCpk beforeCpk afterSigma beforeSigma afterAgainst goal
High-volume product1.151.65139 nm132 nmUnchanged by design
Low-volume product A0.711.17232 nm157 nmGoal met
Low-volume product B0.571.17244 nm176 nm23 nm short of goal

Product A's overlay-y sigma fell from 232 to 157 nanometres, a 75-nanometre improvement that met the goal. Product B fell from 244 to 176, a 68-nanometre improvement that landed 23 nanometres short. I reported it that way. The high-volume product, which the engineers had insisted must not be touched, was untouched: its sigma moved from 139 to 132 nanometres, within noise, and its capability held.

The grouping strategy then became the standard for every run-to-run controller in wafer fabrication at both of Seagate's sites that ran the system. Chemical-mechanical polishing control adopted the identical parameters-database design in production the following quarter, and ion mill control was designed around grouping from the start.

Twenty years later

Still making the drives in today's data centers.

The control system I architected, and the grouping design this project put inside it, are still in production at Seagate two decades on, in four countries. The recording heads for the hard drives in modern data centers are patterned under run-to-run control that traces back to this work. It has outlived the servers it was first installed on, several generations of the tools it controls, and the technology stack it was written in.

The same approach, twenty years apart

The technology changed completely. The method didn't.

In 2005 the stack was Java applets, Enterprise JavaBeans and Oracle. In 2026 it was Python and a local vision model. Put the two projects side by side and the shape of the work is identical, because the shape comes from Six Sigma's DMAIC cycle, not from the tools.

Get in touch

What's on your mind?

A few sentences is enough to start a useful conversation. Just say what your thinking and we'll take it from there.