TechCrunch · AI · 1 ч назад
Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?
Anthropic и OpenAI хотят внедрить независимых оценщиков безопасности в свои лаборатории искусственного интеллекта. Исследователи приветствуют беспрецедентный доступ, но предупреждают, что значимый надзор требует прозрачности, независимости и, в конечном итоге, регулирования.
Подробности
In a lengthy essay published over the weekend, Anthropic CEO Dario Amodei made a proposal that the AI industry would have rejected instantly even a year ago: embed third-party evaluators inside all frontier AI companies, giving them the power to report safety incidents, assess whether AI models are truly aligned, and share their unvarnished findings with the world.
Амодей заявил, что Anthropic обязуется предоставить независимым оценщикам, таким как METR и Redwood Research, беспрецедентный доступ к системам компании. Генеральный директор Сэм Альтман заявил, что OpenAI также будет придерживаться этой практики, сигнализируя о потенциально глубоких изменениях в том, как отрасль работает с внешними исследовательскими группами.
Third-party evaluators who spoke to TechCrunch broadly welcomed the proposal, but said details need to be ironed out — and ideally backed by legislation — if they’re to know whether they will function as truly independent watchdogs or vendors operating on the AI companies’ terms.
Этот более глубокий доступ становится все более важным, поскольку модели начинают лучше распознавать, когда их оценивают, что повышает риск того, что они будут вести себя хорошо во время тестирования, скрывая при этом проблемное поведение. Исследователи говорят, что признаки такого поведения могут быть упущены при тестировании готовой модели, но их можно обнаружить, исследуя, как она вела себя во время обучения.
“AI companies should be able to answer some very basic questions about their training process, such as: Did the AI ever actively try to undermine its own alignment training while it was going through the training?” Alexander Meinke, head of research at Apollo Research, told TechCrunch. “The answer to this should be an unequivocal no, and right now we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public. And we’ve seen from recent incidents that, by default, they will do neither. As embedded evaluators, we could actually check.”