What if using each model to fix the next one makes the risk worse?

The safety story assumes the loop converges. Nothing in the loop guarantees that.

Here is the quiet version of the arms race. Each new model generation gets used to find and patch problems in the next one. On a slide that looks like progress. In practice it is a process with no proof it settles on safety. Every fix can smuggle in a new failure mode that nobody planned for.

That is the same dynamic as the lab-versus-lab sprint, just pointed inward. You are not only racing the other company. You are racing the side effects of your own patches. The system that helps you debug is also the system whose quirks you do not fully see yet.

People talk as if more intelligence applied to the safety problem is automatically more safety. That only holds if the search is aimed well and the errors are the kind that show up in tests. Some errors are quiet. Some look like improvements until they are load-bearing. A loop that keeps scoring "we fixed it" on the dashboard can still be walking toward a worse place.

Hold on. If the models are getting better at finding bugs, why would that not just work?

Because finding a bug is not the same as knowing you did not create two more. Capability can rise faster than understanding. The debugging assistant is a product of the same stack it is inspecting. There is no referee outside the loop. If the process diverges instead of converging, you will not get a clean alarm. You will get a string of green checkmarks and a surprise. That is the meta-risk hiding under the "we use AI to make AI safer" line.