Experts from China, the European Union, the United Kingdom, and the United States have formed the AIxBio Technical Working Group. This group, organized through the AIxBio Global Forum, aims to standardize how AI systems are evaluated for potential misuse in biological contexts.
Why biological evaluations matter
Biological capability evaluations are growing in importance for managing frontier AI risks. These assessments look at what AI systems can do that might make them easier to misuse. Developers use the results to decide on safeguards, access controls, or changes to deployment strategies. However, the lack of consistency in evaluation methods and reporting can lead to misinterpretation of a system's true capabilities.
For example, a system's performance might change depending on how it is configured or tested. If these differences are not clearly reported, they might be wrongly attributed to the model itself. The Working Group will analyze published evaluation methods to create a shared technical basis and standard terminology for evaluating AI systems.
The group's roadmap
The group's tasks include clarifying what these evaluations actually measure and how they relate to the risks they claim to address. They will explore the conditions under which evaluation results can be reliably compared across different systems and organizations. The group will also distinguish what is established by an evaluation from what it implies about the risk of biological misuse.
Evidentiary standards and the thresholds for taking action based on evaluation results will be closely examined. These documents will help developers, governments, and evaluators interpret and use evaluation evidence more consistently.
Open invitation for collaboration
The group is inviting experts and organizations involved in biological capability evaluations to join. They ask participants to contact Helia Samani at [email protected] to contribute or share their work. This effort complements the AIxBio Forum’s existing group on risk assessment and future threats, aiming to fill a critical gap in AI safety protocols.

