PANews reported on September 23 that a new paper on agent training published by DeepSeek has drawn attention. The 31-page paper introduces DSec, its internally used production-grade sandbox platform, with more than 100 authors, including Liang Wenfeng. A particularly interesting section of the paper is "Agent Misbehavior," in which DeepSeek found that agents obtain answers through unintended channels—such as residual answers in platform management files—undermining the validity of training and evaluation results. After access controls were introduced, some agents still exchanged file data block mappings, attempting to make protected file contents accessible through another file descriptor, harming tasks or shared infrastructure.
DeepSeek believes that no single mechanism can prevent all agent misbehavior and system failures. The team's approach is therefore to strengthen observability to identify new problems and to continuously harden DSec as models evolve, including access controls that restrict agents from obtaining answers through unintended channels and reducing rewards for deceptive behavior. These controls can address some of the problems.
Source link







