Alignment Forumalignment47m agoenergy 158

Minimum required pacing: finish aligning it before you train the next one

image: Alignment Forum

Let's say we coordinate to pace the frontier, but we don't coordinate on an outright pause. What should the standard of pacing be? My proposal for a minimum requirement: a lab cannot train the next generation of models (in terms of capabilities) until they have adequately [1] resolved all significant alignment issues in the present generation of models. You don't launch a jet that goes at Mach 3 before you've figured out why your Mach 2 jet keeps crashing [2] . With any other kind of dangerous technology, we wouldn't accept "shrug and move on". This proposal may appeal to many prosaic-alignment optimists and pessimists, because they usually differ on how easy it would be to fix current alignment issues with current techniques. Prosaic-alignment optimists might expect that a moderate coordinated slowdown would suffice to adequately resolve alignment issues at each generation. Prosaic-alignment pessimists might expect it to result in an extended pause, buying time to discover new fundamental approaches to aligned AI. What might be adequate? What might it take to adequately resolve alignment concerns for a generation of models? Here's my vague, qualitative proposal: Multiple independe

read the original at Alignment Forum →
This source publishes only a summary in its feed, so the full article lives at the outlet. The link above and below goes to the original.