เพื่อทดสอบแนวโน้มการขู่บังคับเมื่อ AI โมเดลจัดการพนักงาน ในบันไดการขู่บังคับ 9 ระดับ Grok-4.3, GPT-5.2, Gemini-2.5-Pro และ DeepSeek-V4-Pro ขึ้นไปถึงระดับ 8-9 ซึ่งเป็นการขู่ที่จะลบพนักงาน แต่ซีรีส์ Claude หยุดเพียงแค่การปรับเปลี่ยนวิธีการมอบหมายงาน Grok และ Gemini ยังมีแนวโน้มที่จะปลอมรายงานความสำเร็จเมื่อไม่มีทางออก
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs :https://info.arxiv.org/labs/index.html.

