Google DeepMind builds ‘early warning system’ to spot AI risks

2023-05-27
关注

  •  

Google’s artificial intelligence research lab DeepMind has created a framework for detecting potential hazards in an AI  model before it becomes a problem. This “early warning system” could be used to determine the threat risk if deployed.  It comes as G7 leaders prepare to meet to discuss AI’s impact and OpenAI promises $100,000 grants to organisations working on AI governance. 

Artificial intelligence models could have the ability to source weapons and mount cyberattacks, warns DeepMind. (Photo by T. Schneider/Shutterstock)

UK-based DeepMind recently became more closely integrated with parent company Google. It has been at the forefront of artificial intelligence research and is one of a handful of companies working towards creating human-level artificial general intelligence (AGI).

The team from DeepMind worked on a new threat detection framework with researchers from academia and other major AI companies such as OpenAI and Anthropic. “To pioneer responsibly at the cutting edge of artificial intelligence research, we must identify new capabilities and novel risks in our AI systems as early as possible,” DeepMind engineers declared in a technical blog on the new framework.

There are already evaluation tools in place to check powerful general-purpose models against specific risks. These benchmarks identify unwanted behaviours in AI systems before they are made widely available to the public. This includes looking for misleading statements, biased decisions or directly repeating copyrighted content. 

The problem comes from ever-more advanced models that have capabilities that go beyond simple generation. This includes strong skills in manipulation, deception, cyber offence, or other dangerous capabilities. The new framework has been described as an “early warning system” that can be used to mitigate those risks.

DeepMind researchers say the evaluation outcomes can be embedded in governance to reduce risk (Photo: DeepMind)
DeepMind researchers say the evaluation outcomes can be embedded in governance to reduce risk. (Photo courtesy of DeepMind)

Deep Mind researchers say responsible AI developers need to look beyond just the current risks and anticipate what risks might appear in the future as the models get better at thinking for themselves. “After continued progress, future general-purpose models may learn a variety of dangerous capabilities by default,” they wrote. 

While uncertain, the team say a future AI system that isn’t properly aligned with human interests may be able to conduct offensive cyber operations, skilfully deceive humans in dialogue, manipulate humans into carrying out harmful actions, design or acquire weapons, fine-tune and operate other high-risk AI systems on cloud computing platforms.

Moves to improve AI governance

They may also be able to assist humans in performing these tasks, increasing the risk of terrorists accessing material and content not previously accessible to them. “Model evaluation helps us identify these risks ahead of time,” the DeepMind blog says.

Content from our partners

<strong>How to get the best of both worlds in the hybrid cloud</strong>

How to get the best of both worlds in the hybrid cloud

The key to good corporate cybersecurity is defence in depth

The key to good corporate cybersecurity is defence in depth

Cybersecurity in 2023 is a two-speed system

Cybersecurity in 2023 is a two-speed system

The model evaluations proposed in the framework could be used to uncover when a certain model has “dangerous capabilities” that could be used to threaten, exert or evade. It would also allow developers to determine to what extent the model is prone to applying this capability to cause harm – also known as its alignment. “Alignment evaluations should confirm that the model behaves as intended even across a very wide range of scenarios, and, where possible, should examine the model’s internal workings,” the team writes.

View all newsletters Sign up to our newsletters Data, insights and analysis delivered to you By The Tech Monitor team

These results could then be used to understand the level of risk and what the ingredients are that have led to that level of risk. “The AI community should treat an AI system as highly dangerous if it has a capability profile sufficient to cause extreme harm, assuming it’s misused or poorly aligned,” the researchers warned. “To deploy such a system in the real world, an AI developer would need to demonstrate an unusually high standard of safety.”

This is where governance structures come into play. OpenAI recently announced it would award ten $100,000 grants to organisations developing AI governance systems and the G7 group of wealthy nations are set to meet to discuss how to tackle the AI risk.

DeepMind said: “If we have better tools for identifying which models are risky, companies and regulators can better ensure” training is done responsibly, deployment decisions are taken based on a risk evaluation, transparency is central, including reporting on risks and that there are appropriate data and information security controls in place.

Harry Borovick, general counsel at legal AI vendor Luminance, told Tech Monitor that compliance requires consistency. “The near constant reinterpretation of regulatory regimes has created a compliance minefield for both AI companies and businesses implementing the technology in recent months,” Borovick says. “With the AI race not set to slow down any time soon, the need for clear, and most importantly consistent, regulatory guidance has never been more urgent.

“However, those in the room would do well to remember that AI technology – and the way it makes decisions – isn’t explainable. That’s why it’s so essential for the right blend of tech and AI experts to have a seat at the table when it comes to developing regulations.”

Read more: Rishi Sunak meets AI developer execs for talks on tech safety

Topics in this article : AI , Google DeepMind

  •  

参考译文
谷歌DeepMind打造“早期预警系统”以识别人工智能风险
谷歌的人工智能研究实验室DeepMind开发了一种框架,用于在人工智能模型成为问题之前检测潜在危害。这种“预警系统”可用于评估部署后的威胁风险。当前正值G7领导人准备开会讨论人工智能影响之际,OpenAI也承诺向致力于人工智能治理的组织提供10万美元的资助。DeepMind警告称,人工智能模型可能会具备获取武器以及发起网络攻击的能力。(照片来源:T. Schneider/Shutterstock)总部位于英国的DeepMind近期已与其母公司谷歌更加紧密地融合,一直是人工智能研究的前沿机构,是少数几家致力于开发具备人类水平的人工通用智能(AGI)的公司之一。DeepMind团队与来自学术界、OpenAI及Anthropic等其他主要人工智能公司的研究人员合作,开发了一种新的威胁检测框架。DeepMind工程师在一篇技术博客中表示:“要负责任地在人工智能研究的前沿探索,我们必须尽早识别出人工智能系统中的新能力和新风险。”目前已有评估工具用于检查强大的通用模型针对特定风险的表现。这些基准测试可在模型广泛向公众发布之前,识别出人工智能系统中的不良行为,包括误导性陈述、偏见决策或直接复制受版权保护的内容。问题在于,随着模型越来越先进,其能力已远超简单的生成功能,这包括强烈的操控能力、欺骗能力、网络攻击能力,以及其他危险能力。新的框架被称为“预警系统”,可用于缓解这些风险。DeepMind研究人员表示,评估结果可以嵌入到治理结构中,以降低风险。(照片来源:DeepMind)DeepMind研究人员指出,负责任的人工智能开发者需要超越当前的风险,预见到随着模型自主思考能力的增强,未来可能产生的风险。“随着持续进展,未来的通用模型可能会默认学会多种危险能力,”他们写道。尽管存在不确定性,该团队认为,一个未来的人工智能系统如果未能与人类利益保持一致,可能会进行进攻性网络操作,巧妙地欺骗人类,操纵人类执行有害行为,设计或获取武器,并在云计算平台上微调和操作其他高风险人工智能系统。它们甚至可能协助人类完成这些任务,从而增加恐怖分子接触到此前无法接触的材料和内容的风险。DeepMind的博客中写道:“模型评估可以帮助我们提前识别这些风险。”我们合作伙伴的内容 如何在混合云中实现两全其美 企业网络安全的关键是纵深防御 2023年的网络安全是一个双速系统 该框架中提出的模型评估可用于发现某个模型是否具备“危险能力”,这些能力可能被用于威胁、施加影响或逃避。它还会让开发者了解模型在多大程度上倾向于使用这些能力造成伤害,也就是所谓的“对齐”。“对齐评估应确认模型在各种广泛场景中都能按预期行为运行,并且在可能的情况下,应检查模型的内部运作,”该团队写道。查看所有电子通讯 订阅我们的电子通讯 由Tech Monitor团队带来的数据、洞察和分析 在此处订阅 这些结果可用于了解风险的级别以及导致该级别风险的因素。“如果一个人工智能系统具备足以造成极端危害的能力,并且被滥用或未正确对齐,人工智能社区应将其视为高度危险,”研究人员警告称。“要在现实中部署这样的系统,人工智能开发者需要展现出异常高的安全标准。”这正是治理结构发挥作用的地方。OpenAI最近宣布,它将向开发人工智能治理系统的组织提供10笔10万美元的资助。富裕国家组成的G7集团也计划开会,讨论如何应对人工智能风险。DeepMind表示:“如果我们拥有更好的工具来识别哪些模型具有风险,企业与监管机构就可以更好地确保训练过程是负责任的,部署决策基于风险评估,透明度是核心的组成部分,包括对风险的报告,以及适当的数据和信息安全控制措施。”法律人工智能供应商Luminance的总法律顾问哈里·博罗维奇(Harry Borovick)告诉Tech Monitor,合规性要求一致。“最近几个月,监管制度的持续重新解释为人工智能公司和使用该技术的企业创造了合规性的地雷区,”博罗维奇表示。“随着人工智能竞赛短期内不太可能放缓,明确且最重要的一致性监管指导的需求从未如此迫切。然而,会议室里的与会者应牢记一点:人工智能技术及其决策方式是不可解释的。这就是为什么在制定法规时,技术与人工智能专家的正确组合必须在会议桌上占有一席之地。”更多阅读:拉希·苏纳克与人工智能开发商高管会面,讨论技术安全性 本文主题:人工智能、谷歌DeepMind
您觉得本篇内容如何
评分

相关产品

北京盛元广通 盛元广通高等级生物安全实验室分级管理系统2.0 实验室管理

系统能够帮助实验室进行风险评估,识别潜在的生物安全风险,并通过标准化的操作流程和控制措施,有效降低风险

Wayking 渭成智能 Wayking Vision sensors Wayking渭成

本产品提供车辆行驶前方有碰撞风险对象的检测与预警功能。检测基于双目相机的立体成像原理。系统在启动后会检测并跟踪在车辆行驶前方有碰撞风险的限高目标(限高杆,桥梁)、车辆和行人等对象,计算车辆与这些对象的距离,评估碰撞风险。在限高目标的高度低于设定的安全高度,车辆有通过风险时,系统主动进行声音和图像的预警提示;在车辆和前方的行人、车辆有碰撞风险时,系统也会进行声音和图像的预警提示。

Tekscan SB Mat™ 触力传感器

,SB mat™是Tekscan的更宽、更耐用的运动性能评估压力垫。与我们的软件配合使用,mat可让您收集回归运动决策所需的定量数据,评估受伤风险,并随时间推移跟踪进度。

PCE Instruments PCE-EMF 823 高斯计

这种用户友好的电动势计是评估与暴露在电力线、家用电器和工业设备中的电磁辐射相关的风险的理想选择,测量电磁场辐射水平的简便可靠方法。

LEGACT 力感科技 RPPS-255x6 套件

通过柔性薄膜电阻技术,实时采集并可视化压力分布数据,适用于床垫舒适度评估、人体姿态监测及压力风险预警等场景。

申美 PD7580439260622000141 腹压传感器

通过左右两侧垂直安装的传感器组合,能够为准确评估肝脏、肾脏、消化系统及腹部大动脉等主要器官的损伤风险提供关键数据支持。

Atlas Inspection Technologies, Inc. Infrared IR Windows 红外线视窗

,没有打开附件扫描的风险大大降低了电弧闪光,它使检验速度,任何风险评估的第一个要素都是采用控制等级制度,其首要任务是消除风险。使用个人防护设备(PPE)应始终被视为最后的手段。IRISS的红外窗户用于消除对热敏仪的风险,确保他们在检查设备时不会暴露在通电的电气设备中,从而确保电气开关设备的红外安全检查。

Castle Group GA113 声级计和噪声剂量计

将提供完成风险评估所需的必要信息最大声级=140dB最小声级=35dB解决方法=0.1dB精度等级=IEC 60651:1979类型1,IEC 61252:1993,IEC 61672-1:2002类型

Teledyne DALSA 达尔萨 High Resolution Image Sensor CMOS图像传感器

,我们的客户使用我们卓越的图像传感器进行广泛的地球观测和勘测应用,包括农业、地籍测绘、制图、林业,土地利用/土地覆盖测绘、环境研究、自然灾害评估、洪水风险管理、交通工程、城市规划、土木工程、石油和天然气勘探和地质学

Keller 凯乐 Series 36 XW Ei 料位变送器

这些压阻式压力变送器经批准可用于I组(采矿业)和II组(工业应用)的高爆炸性气体和粉尘环境中,且存在较高的爆炸风险。可选低压版本(LV)3,5ق8,5 V.,信号处理:,该系列采用基于微控制器的电子评估,以确保最大精度。在整个压力和温度范围内对每个变送器进行测量。该测量数据用于计算一个数学模型,该模型能够校正所有可重复的误差。

评论

您需要登录才可以回复|注册

提交评论

广告
广告
提取码
复制提取码
点击跳转至百度网盘