Anthropic publishes new papers showing the latest developments in the training of the AI system involved in improving model alignment. The paper stated that an automated research system had enhanced model performance on 10 matching benchmarks for mismatches, without lowering overall performance.

The automated system completes the study. Take over.

The paper is entitled Automated Researchers Can Reliably Mitigate Aviation Failures. The study was led by Anthropic Fellow Chen Yueh-Han. The system described in the paper will first retrieve existing literature and then propose a training methodology, which will be used to train the model for approximately 30 minutes.

The system will then continue to raise baseline requirements based on results and repeat multiple rounds of testing. Effective methods would be retained and ineffective methods would be eliminated. This process, described by the paper, is close to “checking, proposing, experimenting, sifting” in traditional studies, but implementation is faster and larger.

10 benchmark upgrades

The paper stated that of the 10 alignment benchmarks given, automated systems had been improved in each case and had not resulted in a reduction in overall capacity. This means that not only can the system be modified in response to a particular mismatch, but it can also avoid to some extent the problem of “refining one or the other”.

Anthropic also compared the system with manual researchers. The paper wrote that the best-performing automated alignment research methodology could, on average, exceed the programmes proposed by senior researchers within six hours; and that manual-led research orientation did not lead to stronger results.

Lower cost but still subject to benchmark

The paper also provided a set of cost comparisons: API reasoning costs about $4 per hour for the automated alignment system, while Anthropic pays about $150 per hour for manual researchers.

However, the paper also referred to the apparent limitations of the methodology. The effectiveness of the automated system is provided that the alignment benchmarks themselves accurately reflect the real objectives. If the baseline design is inadequate, the result of system optimization may also deviate from actual needs.

In addition, such systems rely on research documentation and evaluation systems that are maintained on an ongoing basis. In other words, while automated researchers can speed up the testing and screening process, much work remains to be done on benchmarking, updating information and target calibration.