
?¼ÓÄôópc×îÐÂÕ¹ÍûÆÊÎö±¨¸æ?ΪÄãÌṩ¼ÓÄôópc×îÐÂÕ¹ÍûÆÊÎö±¨¸æAPP°²×¿°æÏÂÔØ£¬£¬£¬£¬£¬ÀúÊ·°æ±¾¡¢¾É°æÏÂÔØ£¬£¬£¬£¬£¬Éó²é×îмÓÄôópc×îÐÂÕ¹ÍûÆÊÎö±¨¸æÊÖ»ú°æÏÈÈÝ¡¢Ó¦ÓýØÍ¼¡¢ÍøÓÑ̸ÂÛ£¬£¬£¬£¬£¬Àû±ã¿ì½ÝµÄ½«°²×¿°æ¼ÓÄôópc×îÐÂÕ¹ÍûÆÊÎö±¨¸æÓ¦ÓÃÃâ·ÑÏÂÔØµ½ÊÖ»ú¡£¡£¡£¡£¡£¡£¡£
ÊÇÒ»¿îÂß¼Òâ¼ûÒâÒåÁ¬ÏßÓÎÏ·£¬£¬£¬£¬£¬ÓÎÏ·ÓµÓи»ºñµÄģʽ£¬£¬£¬£¬£¬¶àÑùµÄÍæ·¨£¬£¬£¬£¬£¬Òâ¼ûÒâÒåÒæÖÇ£¬£¬£¬£¬£¬¿ÉÒÔ¸øÓèÍæ¼ÒÐÝÏÐÓéÀÖÌåÑ飬£¬£¬£¬£¬ÒÔΪ²»´íµÄÅóÙ¿ÉÒÔÏÂÔØÊÔÊÔ¡£¡£¡£¡£¡£¡£¡£
ÊÇÒ»¿îÆæÒìµÄ½âÃÕ´³¹ØÓÎÏ·£¬£¬£¬£¬£¬ÔÚÓÎÏ·ÖУ¬£¬£¬£¬£¬Íæ¼Ò±ØÐèѰÕÒè¤Ù¤¹ÝÖÐÒþ²ØµÄÏßË÷ºÍÌáÐÑ£¬£¬£¬£¬£¬Ê¹ÓõÀ¾ßºÍ¼¼ÇÉ×ÊÖú×Ô¼ºÍ»ÆÆÒ»¸ö¸öµÄÃÕÌâ¡£¡£¡£¡£¡£¡£¡£Ã¿Ò»¹ØµÄ½âÃÕÄѶȶ¼Öð½¥ÔöÌí£¬£¬£¬£¬£¬Íæ¼ÒÐèÒª³ä·ÖÑéÕ¹×Ô¼ºµÄÊÓ²ìÁ¦¡¢ÖǻۺÍÄÍÐÄ£¬£¬£¬£¬£¬è¤Ù¤¹ÝÄÚµÄСµÀ¾ß¡¢×°ÊΡ¢Í¼°¸µÈ¶¼Ô̺¬×ŽâÃÕÏßË÷£¡
ÊÇÒ»¿îÕ½ÆåÓÎÏ·£¬£¬£¬£¬£¬ÓÎÏ·ÓµÓи»ºñµÄÎ佫£¬£¬£¬£¬£¬Íæ¼Ò¿ÉÒÔ×÷ÓýÎ佫£¬£¬£¬£¬£¬¿ËµÐÖÆÊ¤£¬£¬£¬£¬£¬¾çÇ鸻ºñ£¬£¬£¬£¬£¬Íæ·¨¶àÑù£¬£¬£¬£¬£¬Ï²»¶Õ½ÆåÓÎÏ·µÄÅóÙ½Ó´ýǰÀ´ÏÂÔØË¬Íæ¡£¡£¡£¡£¡£¡£¡£
ÊÇÒ»¿î¼ÓÈë¸Ð½ÏÇ¿µÄ¶ùͯÉúÑÄÓÎÏ·£¡ÕâÊÇ¿îÏñÒ»±¾ÊéÒ»ÑùµÄÓÎÏ·£¬£¬£¬£¬£¬¼ÓÈëÄãµÄÅóÙÁÚÈ˼ң¬£¬£¬£¬£¬ÕâÀïËæ´¦¶¼ÊÇÓѺõÄÅóÙÃæÄ¿£¬£¬£¬£¬£¬ÔÚÕâÀïÄã¿ÉÒÔºÍһЩС¶¯ÎïÁÚÈËÒ»ÆðÍæË££¬£¬£¬£¬£¬ÔÆÔÆÓÐȤ£¬£¬£¬£¬£¬ÔÆÔÆÎüÒýÈË£¬£¬£¬£¬£¬¿ìÀ´ÏÂÔØÌåÑé°É£¡
ÊÇÒ»¿îÊ®·Ö´Ì¼¤µÄÈü³µ¾ºËÙÓÎÏ·£¬£¬£¬£¬£¬ÕæÊµÄ£Äâ³öÁËÕæÊµµÄ¼ÝÊ»³¡¾°£¬£¬£¬£¬£¬»Öʼ«Îª¸ßÇåϸÄ壬£¬£¬£¬£¬ÕæÊµ¸ÐÂúÂú£¬£¬£¬£¬£¬´ø¸øÍæ¼ÒËÙÂÊÓ뼤Çé°ãµÄÓÎÏ·¼ÝÊ»ÌåÑé¡£¡£¡£¡£¡£¡£¡£
ÊÇÒ»¿îÍæ·¨ºÜÓÐȤ²¢ÇÒÌåÑé¸ÐÒ²ºÜ½âѹµÄÐÝÏÐÄ£Ä⿪³µÐ¡ÓÎÓΣ¬£¬£¬£¬£¬Õâ¿îÓÎÏ·ÐèÒª½«×Ô¼ºµÄ³µÁ¾ÍêÉÆµÄ´ÓÍ£³µÎ»ÉÏ¿ª×ߣ¬£¬£¬£¬£¬¶ÔÕâ¿îÓÎÏ·¸ÐÐËȤµÄÍæ¼Ò¾Í¿ìÀ´ÏÂÔØ°É£¡
×ªÔØ×Ô£ºminimax Ï¡Óî¼ÓÄôópc×îÐÂÕ¹ÍûÆÊÎö±¨¸æ¿Æ¼¼
Ëæ×Å minimax m2.5 µÄÐû²¼²¢ÔÚÉçÇøÒý·¢ÈÈÁÒ»ØÉù£¬£¬£¬£¬£¬ºÜÐË·ÜÄܽè´Ëʱ»ú£¬£¬£¬£¬£¬·ÖÏíÔÚÄ£×ÓѵÁ·±³ºó¹ØÓÚ agent rl ϵͳµÄһЩ˼Ë÷¡£¡£¡£¡£¡£¡£¡£
ÔÚ´ó¹æÄ£¡¢ÖØ´óµÄÕæÊµÌìϳ¡¾°ÖÐÅÜ rl ʱ£¬£¬£¬£¬£¬Ê¼ÖÕÃæÁÙÒ»¸ö½¹µãÄÑÌ⣺ÔõÑùÔÚϵͳÍÌÍÂÁ¿¡¢ÑµÁ·ÎȹÌÐÔÓë agent ÎÞаÐÔÕâÈýÕßÖ®¼äÈ¡µÃƽºâ¡£¡£¡£¡£¡£¡£¡£ÎªÏàʶ¾öÕâ¸öÎÊÌ⣬£¬£¬£¬£¬ÎÒÃÇÉè¼ÆÁËÒ»¸öÒì²½µÄÔÉú agent rl ϵͳ¡ª¡ª forge¡£¡£¡£¡£¡£¡£¡£ÔÚ forge ÖУ¬£¬£¬£¬£¬ÎÒÃÇͨ¹ýʵÏÖ±ê×¼»¯µÄ agent-llm ½»»¥ÐÒ飬£¬£¬£¬£¬Ö§³ÖÁ˶Ôí§Òâ agent ½ÅÊּܾÙÐÐѵÁ·£¬£¬£¬£¬£¬²¢ÇÒͨ¹ý¼«ÖµĹ¤³ÌÓÅ»¯ºÍÎȹ̵ÄËã·¨Óë½±ÀøÉè¼Æ£¬£¬£¬£¬£¬ÊµÏÖÁ˳¬´ó¹æÄ£µÄÇ¿»¯Ñ§Ï°¡£¡£¡£¡£¡£¡£¡£
ÔÚÃæÁÙÊýÊ®Íò¸öÕæÊµµÄ agent ½ÅÊּܺÍÇéÐÎÒÔ¼° 200k µÄÉÏÏÂÎij¤¶Èʱ£¬£¬£¬£¬£¬PTÊÓѶ(ÖйúÇø)¹ÙÍø rl ϵͳ×öµ½ÁËÌìÌì°ÙÍò¼¶Ñù±¾Á¿µÄÍÌÍ£¬£¬£¬£¬£¬²¢ÊµÏÖÒ»Á¬ÎÈ¹ÌµÄ reward ÉÏÕǺÍÕæÊµµÄÄ£×ÓÄÜÁ¦ÌáÉý£¬£¬£¬£¬£¬²¢×îÖÕ×÷ÓýÁË minimax m2.5 Ä£×ÓµÄÐÔÄÜÍ»ÆÆ¡£¡£¡£¡£¡£¡£¡£
ÎÊÌ⽨ģ
ÔÚÉîÈë̽Ìּܹ¹Éè¼ÆÖ®Ç°£¬£¬£¬£¬£¬ÎÒÃÇÊ×ÏȽ« agent Ç¿»¯Ñ§Ï°ÏµÍ³µÄÓÅ»¯Ä¿µÄÐÎʽ»¯Îª¡°×î´ó»¯ÓÐÓÃѵÁ·ÊÕÒæ£¨j£©¡±£º
![]()
ÆäÖУ¬£¬£¬£¬£¬throughput ÊÇָÿÃë´¦Öóͷ£µÄÔʼ token ÊýÄ¿£¬£¬£¬£¬£¬ÆäÖ÷ÒªÊÜ rl ϵͳÖеÄËIJ¿·Ö¿ØÖÆ£º rollout¡¢training¡¢data processing ºÍ i/o¡£¡£¡£¡£¡£¡£¡£sample efficiency ÔòÊÇָÿ¸öÑù±¾´øÀ´µÄƽ¾ùÐÔÄÜÌáÉý£¬£¬£¬£¬£¬ÓÉÊý¾ÝÂþÑÜ¡¢Êý¾ÝÖÊÁ¿¡¢Ë㷨ЧÂÊÒÔ¼° offpolicy ˮƽ¾öÒé¡£¡£¡£¡£¡£¡£¡£¶øÎȹÌÐÔºÍÊÕÁ²ÐÔÔòÄܹ»»ùÓÚѵÁ·Àú³ÌÖмà²âÖ¸±êÀ´Åжϡ£¡£¡£¡£¡£¡£¡£
ҪʵÏÖ£¨j£©µÄ×î´ó»¯£¬£¬£¬£¬£¬ÎÒÃÇÐèҪսʤÒÔÏÂÈýÀàÌôÕ½£º
Ä¿½ñ³£¼ûµÄ rl ¿ò¼ÜºÍ·¶Ê½¶Ô agent µÄÖØÆ¯ºóÏÞÖÆºÜ´ó£¬£¬£¬£¬£¬Ö÷ÒªÌåÏÖÔÚ£º
agent ×ÔÓɶÈÊÜÏÞ£º½« agent ÊÓΪ°×ºÐ¾ÍÒªÇóÔÚ agent ºÍ rl framework Ö®¼ä¹²ÏíºÍת´ï״̬¡£¡£¡£¡£¡£¡£¡£ÕâÖÖÉè¼ÆÄÑÒÔ¶ÔÖØ´óµÄ agent ¼Ü¹¹£¨È綯̬ÉÏÏÂÎÄÖÎÀí¡¢multi-agent rl µÈ£©¾ÙÐн¨Ä££¬£¬£¬£¬£¬µ¼ÖÂÄ£×ÓÄÜÁ¦ÎÞ·¨ÔÚÖØ´óµÄºÚºÐ agent ÉÏÓÐÓ÷º»¯¡£¡£¡£¡£¡£¡£¡£
tokenÒ»ÖÂÐÔÎÊÌ⣺ÏÖÓÐµÄ tito£¨token-in-token-out£©Ä£Ê½ÆÈʹ agent Óëµ×²ãµÄ tokenizer Âß¼Éî¶ÈñîºÏ¡£¡£¡£¡£¡£¡£¡£ÔÚÖØ´óµÄÉÏÏÂÎÄÖÎÀí»úÖÆÏ£¬£¬£¬£¬£¬ÒªÏëά³Ö agent ºÍ rl Ö®¼äµÄÑÏ¿áÒ»ÖÂÐÔ£¬£¬£¬£¬£¬Æä¹¤³Ì±¾Ç®ÊǺÜÊÇ´óµÄ¡£¡£¡£¡£¡£¡£¡£
rollout µÄÍê³Éʱ¼ä±£´æ¼«´óµÄ·½²î¡ª¡ª¶ÌÔò¼¸Ã볤ÔòÊýСʱ¡£¡£¡£¡£¡£¡£¡£Õâ´øÀ´ÁËÒ»¸öÒì³ÌÐò¶ÈÎÊÌ⣺
ÑµÍÆÒì³ÌÐò¶ÈÂß¼£ºÅܹýÒì²½ rl µÄͬÑâ¶¼ÖªµÀ£¬£¬£¬£¬£¬ÔÚ mfu ºÍ rl Ëã·¨ÎȹÌÐÔÖ®¼äȨºâÊǺÜÊÇÖØ´óµÄ¡£¡£¡£¡£¡£¡£¡£ÑÏ¿áµÄ fifo£¨first in first out£©/ͬ³ÌÐò¶È»á±»ÓÚ³¤Î²Ñù±¾ block£»£»£»£»£»£»£»£»¶ø greedy/fffo£¨first finish first out£©ËäÈ»×î´ó»¯ÁËÍÌÍÂÁ¿£¬£¬£¬£¬£¬È´´øÀ´Á˲»¿É¿ØµÄ distribution shift£¬£¬£¬£¬£¬¼«Ò×µ¼Ö rl ÖÐ;±Àµô¡£¡£¡£¡£¡£¡£¡£
ǰ׺ÈßÓࣺÔÚ¶àÂÖ agent ÇëÇóºÍ group-level µÄ rollout ÖУ¬£¬£¬£¬£¬tokenizer µÄ encode-decode ·×ÆçÖÂÐÔºÍÉÏÏÂÎÄÖÎÀí»úÖÆ£¬£¬£¬£¬£¬»áµ¼ÖÂÇëÇó¼ä¹²ÏíÁË´ó×ÚµÄǰ׺£¬£¬£¬£¬£¬ÕâÖÖÈßÓàÔÚѵÁ·Ê±´úÔì³ÉÁËÖØ´óµÄÅÌËãÆÌÕÅ¡£¡£¡£¡£¡£¡£¡£
Ï£º±½±ÀøÎÊÌâ£ºÖØ´óµÄ agent ʹÃüµÄ trajectory ͨ³£°üÀ¨³¤´ïÊýǧ²½£¬£¬£¬£¬£¬Ê¹µÃ»ùÓÚÏ£º±½±ÀøµÄ credit assignment ÔÚÊýѧÉϺÜÊDz»Îȹ̡£¡£¡£¡£¡£¡£¡£ÕâÖÖÏ£º±ÐÔµ¼Ö»ر¨ÅÌËãÖеÄÐÅÔë±È¼«µÍ£¬£¬£¬£¬£¬ÒýÆð¸ßÌݶȷ½²î£¬£¬£¬£¬£¬ÆÆËðÁË´ó¹æÄ£Ä£×ÓѵÁ·µÄÎȹÌÐÔ¡£¡£¡£¡£¡£¡£¡£
long cot µÄ¸ºÃæÓ°Ï죺ÔÚ r1 ³öÀ´Ö®ºó¸÷È赀 rl ¶¼ºÜ¹Ø×¢ response length µÄÔöÌí¡£¡£¡£¡£¡£¡£¡£µ«ÔÚÕæÊµµÄ agent ³¡¾°ÖУ¬£¬£¬£¬£¬Óû§×Åʵ¶ÔÖ´ÐÐʱ¼äºÜÊǹØ×¢£¬£¬£¬£¬£¬ÈôÊDz»¼ÓÒÔÏÞÖÆ¿ÉÄܻᵼÖÂѵ³öÀ´µÄÄ£×ÓËäȻˢ°ñºÜÇ¿£¬£¬£¬£¬£¬µ«Óû§ÌåÑéºÜ²î¡£¡£¡£¡£¡£¡£¡£
ϵͳ¼Ü¹¹Óëagent rl·¶Ê½
rl ϵͳÉè¼Æ
ΪÁËʵÏÖÕæÕý¿ÉÀ©Õ¹µÄ¼Ü¹¹£¬£¬£¬£¬£¬ÎÒÃDz»ÔÙ¾ÖÏÞÓÚÏêϸµÄ agent£¬£¬£¬£¬£¬¶øÊÇתÏòÁËͨÓõÄÁýͳ²ãÉè¼Æ£¬£¬£¬£¬£¬½« agent µÄÖ´ÐÐÂß¼Óëµ×²ãµÄÑµÍÆÒýÇæ³¹µ×½âñî¡£¡£¡£¡£¡£¡£¡£PTÊÓѶ(ÖйúÇø)¹ÙÍø rl ϵͳÓÉ 3 ¸ö½¹µãÄ£¿£¿£¿£¿£¿£¿é×é³É£º
![]()
1.agent£º¸Ã²ãÁýͳÁËͨÓà agent£¨º¸Ç°×ºÐºÍºÚºÐ¼Ü¹¹£©¼°ÆäÔËÐÐÇéÐΡ£¡£¡£¡£¡£¡£¡£ËüÈÏÕæÐµ÷ÇéÐν»»¥£¬£¬£¬£¬£¬Ê¹ agent ³ÉΪһ¸ö´¿´âµÄ trajectory producer¡£¡£¡£¡£¡£¡£¡£Í¨¹ý½«ÇéÐν»»¥Óë llm generation ½âñ£¬£¬£¬£¬agent ¿ÉÒÔרעÓÚ½¹µãÓªÒµÂß¼£¨Èç context management ºÍÖØ´óµÄÇéÐν»»¥µÈ£©£¬£¬£¬£¬£¬¶øÎÞÐèÌåÌùµ×²ãµÄѵÁ·ºÍÍÆÀíϸ½Ú¡£¡£¡£¡£¡£¡£¡£
2.ÖÐÐļþÁýͳ²ã£º×÷ΪÇÅÁº£¬£¬£¬£¬£¬¸Ã²ãÔÚÎïÀíÉϽ« agent ²àÓëѵÁ·/ÍÆÀíÒýÇæ¸ôÀë¡£¡£¡£¡£¡£¡£¡£
gateway server£º³äµ±±ê×¼»¯Í¨Ñ¶Íø¹Ø£¬£¬£¬£¬£¬´¦Öóͷ£ agent Óë llm Ö®¼äµÄ½»»¥ÇëÇ󡣡£¡£¡£¡£¡£¡£Í¨¹ýͨÓñê×¼ÐÒ飬£¬£¬£¬£¬ËüÓÐÓõؽ«µ×²ãÄ£×ÓµÄÖØ´óÐÔÓë agent µÄ¸ß²ãÐÐΪÂß¼¸ôÍÑÀëÀ´¡£¡£¡£¡£¡£¡£¡£
data pool£º×÷ΪÂþÑÜʽÊý¾Ý´æ´¢£¬£¬£¬£¬£¬Òì²½ÍøÂç trajectory ºÍ process signal¡£¡£¡£¡£¡£¡£¡£Ëü³äµ±ÌìÉúºÍѵÁ·½âñîµÄ»º³åÇø£¬£¬£¬£¬£¬ÔÊÐíÎÞаµÄÊý¾Ý´¦Öóͷ£ºÍÅú´¦Öóͷ£Õ½ÂÔ¡£¡£¡£¡£¡£¡£¡£
3.ѵÁ·ÓëÍÆÀíÒýÇæ£º
rollout engine£º×¨ÓÃÓÚ¸ßÍÌÍÂÁ¿ token ÌìÉú£¬£¬£¬£¬£¬ÏìÓ¦ agent µÄÌìÉúÇëÇ󡣡£¡£¡£¡£¡£¡£
train engine£ºÍ¨¹ý scheduler ´Ó data pool ÖÐ fetch Êý¾Ý£¬£¬£¬£¬£¬¸üРagent model£¬£¬£¬£¬£¬²¢Óë²ÉÑùÒýÇæ¼á³Öͬ²½£¬£¬£¬£¬£¬È·±£ agent ʹÓÃ×îеÄÕ½ÂÔÂþÑܾÙÐÐ̽Ë÷¡£¡£¡£¡£¡£¡£¡£
ÎÒÃÇÔÚÀëÏ߯À¹ÀÖз¢Ã÷£¬£¬£¬£¬£¬²î±ð agent ½ÅÊּܻᵼÖÂÏÔÖøµÄÐÔÄÜÎó²î¡£¡£¡£¡£¡£¡£¡£½èÖú¸ÃÄ£¿£¿£¿£¿£¿£¿é»¯Éè¼Æ£¬£¬£¬£¬£¬ÎÒÃÇÔÚÎÞÐèÐÞ¸Ä agent ÄÚ²¿´úÂëµÄÇéÐÎÏ£¬£¬£¬£¬£¬Ê¹Óôó×ÚµÄ agent ¿ò¼Ü¾ÙÐÐÁËѵÁ·¡£¡£¡£¡£¡£¡£¡£ÕâÖÖ¡°ÒýÇæÓë agent ÍêÈ«½âñµÄ¼Ü¹¹È·±£ÁËÄ£×ÓÄÜÔÚÖÖÖÖÇéÐÎÖзº»¯£¬£¬£¬£¬£¬ÏÖÔÚÎÒÃÇÒѼ¯³ÉÁËÊý°ÙÖÖ¿ò¼ÜºÍÊýǧÖÖ²î±ðµÄ¹¤¾ßŲÓÃÃûÌᣡ£¡£¡£¡£¡£¡£
¹ØÓڰ׺Рagent£¬£¬£¬£¬£¬ÎÒÃÇ¿ÉÒÔͨ¹ý³ä·ÖµÄ½ÅÊÖ¼ÜÉè¼ÆºÍÔö¹ã£¬£¬£¬£¬£¬ÒÔÖ±½ÓÊÓ²ìºÍÓÅ»¯Ä£×ÓÔÚÌØ¶¨ÀàÐÍ agent ÉϵÄÌåÏÖ¡£¡£¡£¡£¡£¡£¡£ÔÚ m2.5 ÖУ¬£¬£¬£¬£¬ÎÒÃÇÌØÊâÓÅ»¯ÁËÒÑÍùÄ£×ÓÔÚ´øÉÏÏÂÎÄÖÎÀíµÄ³¤³ÌʹÃü£¨Èç deepsearch£©ÖзºÆðµÄһЩÎÊÌ⣺
ÉÏÏÂÎij¡¾°ÐÔÄÜÍË»¯£ºËæ×Ž»»¥ÂÖ´ÎÔöÌí£¬£¬£¬£¬£¬ÖÐÐÄÍÆÀíºÍÈßÓàÊÓ²ìµÄ»ýÀۻᱬ·¢¡°×¢ÖØÁ¦Ï¡ÊÍ¡±¡£¡£¡£¡£¡£¡£¡£ÕâÖÖÔëÉù»áµ¼ÖÂÄ£×ÓÔÚ¾ø¶ÔÉÏÏÂÎÄ´°¿ÚÄÚ¶ÔÒªº¦ÐÅϢʧȥ½¹µã¡£¡£¡£¡£¡£¡£¡£
ÑµÍÆ·×ÆçÖ£ºËäÈ»ÉÏÏÂÎÄÖÎÀí¿ÉÒÔÑÓÉì½»»¥ÖÜÆÚ£¬£¬£¬£¬£¬ÌáÉý agent ÔÚ³¤ÉÏÏÂÎij¡¾°µÄÌåÏÖ£¬£¬£¬£¬£¬µ«½öÔÚÍÆÀíʱʹÓûáÓÉÓÚÆ«Àë rl ѵÁ·µÄÊý¾ÝÂþÑÜ£¬£¬£¬£¬£¬ÆÈʹģ×ÓÔÚÍÆÀíʱ±»ÆÈ½ÓÊÜÉÏÏÂÎıäǨ£¬£¬£¬£¬£¬´¦Öóͷ£²»³£¼ûµÄ³¤ÏÂÎÄ£¬£¬£¬£¬£¬´Ó¶øÓ°ÏìÄ£×ÓÌåÏÖ¡£¡£¡£¡£¡£¡£¡£
ΪÏàʶ¾öÕâЩÎÊÌ⣬£¬£¬£¬£¬ÎÒÃǽ«ÉÏÏÂÎÄÖÎÀí£¨context management£¬£¬£¬£¬£¬ cm£©»úÖÆÖ±½ÓÕûºÏµ½ rl ½»»¥Ñ»·ÖУ¬£¬£¬£¬£¬½«ÆäÊÓΪÇý¶¯×´Ì¬×ª»»µÄ¹¦Ð§ÐÔÐж¯£º
cm Çý¶¯µÄ״̬ת»»£ºÎÒÃǽ« cm ½¨Ä£Îª agent action£¬£¬£¬£¬£¬¶øÉÏÏÂÎıäǨÔòÔ̺¬ÔÚÇéÐ뵀 dynamics ÖС£¡£¡£¡£¡£¡£¡£×´Ì¬´Ó s£¨t£©µ½ s£¨t+1£©µÄת»»Òþʽ°üÀ¨ÁËÉÏÏÂÎÄÇл»µÄÂß¼£¬£¬£¬£¬£¬½«ÉÏÏÂÎÄ˳Ӧ°üÀ¨ÔÚÁËÄ£×ÓµÄѵÁ·Ä¿µÄÖС£¡£¡£¡£¡£¡£¡£
×ÔË³Ó¦ÍÆÀíģʽ£ºÍ¨¹ýÔÚ´Ë¿ò¼ÜÄÚÓÅ»¯Õ½ÂÔ ¦Ð£¨¦È£©£¬£¬£¬£¬£¬Ä£×Óѧ»áÁËÄÚ»¯ÂþÑÜÆ«ÒÆ£¬£¬£¬£¬£¬Ó¿ÏÖ³öÓÅÏȹØ×¢ state-critical token µÄ³°ôÍÆÀíģʽ¡£¡£¡£¡£¡£¡£¡£
¸ÐÖªÉÏÏÂÎÄÖÎÀíÕ½ÂÔ£ºÔÚ¸ÃÕ½ÂÔÏ£¬£¬£¬£¬£¬Ä£×ÓÔÚ rl ÌìÉúÀú³ÌÖоÍÐèҪѧ»áÔ¤¼û¿ÉÄܵÄÉÏÏÂÎÄÖÎÀíºÍ¸Ä±ä£¬£¬£¬£¬£¬Ä£×Óͨ¹ý×Ô¶¯±£´æÓëÄ¿µÄʹÃüÏà¹ØµÄÐÅÏ¢ºÍïÔÌÎÞ¹ØÉÏÏÂÎÄÐÅÏ¢£¬£¬£¬£¬£¬´ó·ùÌáÉýÁËÔÚ context-management agent ϵÄÐÔÄÜ¡£¡£¡£¡£¡£¡£¡£
Ðí¶àÓû§µÄÕæÕýÔÚÓÃµÄ agent ÏÖʵÉÏÊDZÕÔ´µÄ£¬£¬£¬£¬£¬ÎÒÃÇÍêÈ«ÎÞ·¨¸ÐÖªÄÚ²¿µÄ agent loop Âß¼¡£¡£¡£¡£¡£¡£¡£ÎªÁËÈ·±£Ä£×ÓÔÚ²»Í¸Ã÷¼Ü¹¹ÉÏÒ²ÄܶԽÅÊÖ¼ÜÕë¶ÔÐÔÓÅ»¯£¬£¬£¬£¬£¬ÎÒÃǽÓÄÉÁËÒÔϼƻ®£º
·ÇÇÖÈëʽ¼¯³É£ºforge ²»¸ÐÖª agent ÄÚ²¿µÄʵÏÖϸ½Ú£¬£¬£¬£¬£¬ÄÚ²¿Ö»ÐèÒª½«ÇëÇó´òµ½ rl ЧÀ굀 gateway£¬£¬£¬£¬£¬¿ò¼ÜÄÚ²¿¼´¿É¾ÙÐÐÊý¾ÝÍøÂçºÍѵÁ·£¬£¬£¬£¬£¬Òò´ËÔÚÏÖʵ rl ѵÁ·Ê±¿ÉÒÔ¼æÈÝí§ÒâÉÏÏÂÎIJÙ×÷£¨ÈçÓ°ÏóѹËõ¡¢ÀúÊ·ÖØÐ´£©£¬£¬£¬£¬£¬í§ÒâÄÚ²¿µÄ agent loop£¨ÀýÈç deep think¡¢multi-agent µÈµÈ£©¡£¡£¡£¡£¡£¡£¡£
¶à¿ò¼Ü·º»¯£ºÍ¨¹ý½«ÑµÁ·Ñ»·Óë agent ÄÚ²¿×´Ì¬½âñ£¬£¬£¬£¬minimax m2.5 ÆÕ±éÊÊÅä´ó×ںںРagent¡ª¡ªÎÞÂÛÊÇÒÔɳºÐ+mcp ÇéÐÎΪÖ÷µÄ´úÂë agent£¨ÀýÈçÎÒÃǽ« opencode agent Ö±½ÓÊÓΪһ¸öºÚºÐ agent À´ÑµÁ·£©£¬£¬£¬£¬£¬ÕÕ¾ÉʹÓü¤½øÉÏÏÂÎÄËõ¼õÕ½Â﵀ agent£¨Èç truncate bc£©¡£¡£¡£¡£¡£¡£¡£ÊµÑéÅú×¢£¬£¬£¬£¬£¬¸ÃÒªÁìÔÚÍêÈ«²»Í¸Ã÷µÄºÚºÐϵͳÉÏÒÀÈ»ÄÜ´øÀ´Îȹ̵ÄÌáÉý¡£¡£¡£¡£¡£¡£¡£
![]()
¹¤³ÌÓÅ»¯
ΪÏàʶ¾öÍÌÍÂÁ¿ÓëÊý¾ÝÂþÑÜÒ»ÖÂÐÔÖ®¼äµÄ³åÍ»£¬£¬£¬£¬£¬ÎÒÃÇÌá³öÁË windowed fifo µ÷ÀíÕ½ÂÔ¡£¡£¡£¡£¡£¡£¡£¸ÃÕ½ÂÔ½éÓÚ fifo ºÍ greedy Ö®¼ä£¬£¬£¬£¬£¬¼´¿ÉÒÔ°ü¹ÜϵͳµÄÍÌÍ£¬£¬£¬£¬£¬Ò²¿ØÖÆÁËÑù±¾µÄ off-policyness¡£¡£¡£¡£¡£¡£¡£
¼ÙÉèÄ¿½ñµÖ´ïÁË×î´óµÄÌìÉú²¢·¢Á¿£¨Èç n = 8192£©£¬£¬£¬£¬£¬ÌìÉúÐÐÁÐΪ q£¬£¬£¬£¬£¬Ä¿½ñÍ·²¿Î»ÓÚË÷Òý h¡£¡£¡£¡£¡£¡£¡£ÑµÁ·µ÷ÀíÆ÷ÊÜÏÞÓÚÒ»¸ö¾ÞϸΪw£¨Èç w=4096£©µÄ¿É¼û´°¿Ú£º
ÊÜÏ޿ɼûÐÔ£ºµ÷ÀíÆ÷Ö»ÄÜ´Ó¹æÄ£ÄÚ»ñÈ¡ÒÑÍê³ÉµÄ¹ì¼£¡£¡£¡£¡£¡£¡£¡£
¾Ö²¿Ì°ÐÄ£¨´°¿ÚÄÚ£©£ºÔÚÔ˶¯´°¿ÚÄÚ£¬£¬£¬£¬£¬µ÷ÀíÆ÷¿ÉÁ¬Ã¦ÌáÈ¡ÈκÎÒÑÍê³É¹ì¼££¬£¬£¬£¬£¬×èÖ¹Á˶ÓÍ·ÛÕ±Õ£¨hol£©£¬£¬£¬£¬£¬¿ìËÙʹÃüÎÞÐèÆÚ´ýÍ·²¿Ê¹ÃüÍê³É¡£¡£¡£¡£¡£¡£¡£
È«¾ÖÑÏ¿áÛÕ±Õ£¨´°¿ÚÍ⣩£º×ÝÈ»Ë÷ÒýΪ h+w+k µÄʹÃüÒÑÍê³É£¬£¬£¬£¬£¬µ÷ÀíÆ÷Ҳեȡ»ñÈ¡Ëü¡£¡£¡£¡£¡£¡£¡£
Ô¼ÊøÍÆ½ø£ºÖ»Óе±Í·²¿µÄʹÃü±»ÏûºÄʱ£¬£¬£¬£¬£¬´°Ì¸·æÏòǰ»¬¶¯£¨h¡úh+1£©¡£¡£¡£¡£¡£¡£¡£ÕâÆÈʹµ÷ÀíÆ÷±ØÐèÆÚ´ýÄ¿½ñ´°¿ÚÄڵġ°³¤ÖÜÆÚÂäÎéʹÃü¡±£¬£¬£¬£¬£¬±ÜÃâѵÁ·ÂþÑÜÏò¡°¿ì¶ø¼òÆÓ¡±µÄÑù±¾ÑÏÖØÆ«ÒÆ¡£¡£¡£¡£¡£¡£¡£
![]()
agent µÄ¶àÂÖÇëÇó¼ä±£´æºÜ¸ßµÄÉÏÏÂÎÄÇ°×ºÖØºÏ¶È£¬£¬£¬£¬£¬¹Å°åÒªÁ콫ÿ¸öÇëÇóÊÓΪ×ÔÁ¦Ñù±¾£¬£¬£¬£¬£¬Öظ´ÅÌË㹫¹²Ç°×º£¬£¬£¬£¬£¬ÆÌÕÅÁË´ó×ÚµÄÑ·üçãÁ¦¡£¡£¡£¡£¡£¡£¡£
ÎÒÃÇÌá³öÁË prefix tree merging ¼Æ»®£¬£¬£¬£¬£¬½«ÑµÁ·Ñù±¾´Ó¡°ÏßÐÔÐòÁС±Öع¹Îª¡°Ê÷Ðνṹ¡±£¬£¬£¬£¬£¬ÏÂÃæÊÇÏêϸµÄÊý¾Ý´¦Öóͷ£ºÍѵÁ·Õ½ÂÔ£º
Ö»Òª¹²Ïí»ù´¡Ç°×º£¬£¬£¬£¬£¬completions ¾ÍÄÜÔÚÑù±¾¼¶±ðºÏ²¢µ½Ò»¿Ãǰ׺Ê÷ÖУ¨×ÝÈ»ºóÐøÏìÓ¦»ò²ÉÑù·ÖÖ§²î±ð£©¡£¡£¡£¡£¡£¡£¡£
ͨ¹ýʹÓà attention mask ÔÓÈç magi attention£©ÌåÏÖ²î±ð branch Ö®¼äµÄÒÀÀµ¹ØÏµ£¬£¬£¬£¬£¬¿ÉÒÔ°ü¹ÜǰÏòÅÌËãÔÚÊýѧÉÏÓë naive ¼Æ»®ÍêȫһÖ£¬£¬£¬£¬£¬ÔÚÅÌËã loss ʱ£¬£¬£¬£¬£¬ÎÒÃÇ»á°Ñǰ׺Ê÷ unmerge ΪÐòÁеÄÃûÌ㬣¬£¬£¬£¬²»Ó°ÏìºóÐøµÄ loss ÅÌËãºÍÖ¸±êͳ¼Æ¡£¡£¡£¡£¡£¡£¡£
¸Ã¼Æ»®Ïû³ýÁËÈßÓàµÄǰ׺£¬£¬£¬£¬£¬Ïà±ÈÓÚ naive ¼Æ»®ÊµÏÖÁËÔ¼ 40 ±¶µÄѵÁ·¼ÓËÙ£¬£¬£¬£¬£¬ÇÒÏÔÖø½µµÍÁËÏԴ濪Ïú¡£¡£¡£¡£¡£¡£¡£
![]()
ÒýÈëÒì²½ rl Ö®ºóËäÈ» rollout ½×¶ÎËãÁ¦Õ¼±È½µµÍµ½ÁË 60% ×óÓÒ£¬£¬£¬£¬£¬µ«ÍÆÀí×Ô¼ºÉÐÓкܴóÓÅ»¯¿Õ¼ä£¬£¬£¬£¬£¬ÎÒÃÇͨ¹ýÏÂÃæµÄ¼¸ÏîÓÅ»¯À´¼ÓËÙ llm ÍÆÀí£º
dynamic mtp£ºÊ×ÏÈÎÒÃÇÒýÈë mtp ¾ÙÐÐÍÆÀí¼ÓËÙ£¬£¬£¬£¬£¬Í¬Ê±ÎªÁ˰ü¹ÜѵÁ·Àú³ÌÖÐά³Ö draft model µÄ¸ß½ÓÊÜÂÊ£¬£¬£¬£¬£¬ÎÒÃÇͨ¹ý top-k kl loss ÔÚ rl Àú³ÌÖÐÒ»Á¬ÑµÁ· detached mtp head£¬£¬£¬£¬£¬Óë rl policy ¼á³Ö¶ÔÆë¡£¡£¡£¡£¡£¡£¡£
rollout ²àµÄ pd ÊèÉ¢£ºpd ÊèÉ¢¿ÉÒÔÏû³ý moe µ÷ÀíÖÐµÄ pd ×ÌÈÅ£¬£¬£¬£¬£¬ÎªÃ¿¸öʵÀýÌṩ×ÔÁ¦µÄ²¢ÐкÍÌìÉúÕ½ÂÔ£¬£¬£¬£¬£¬ÔÚ×î´ó»¯ÍÌÍÂÁ¿µÄͬʱÓÅ»¯³¤Î²Ñù±¾µÄÑÓ³Ù£¬£¬£¬£¬£¬±ÜÃ⼫¶ËÑù±¾ÛÕ±Õ fifo scheduler£¬£¬£¬£¬£¬²¢´øÀ´½Ï¸ßµÄ offpolicy¡£¡£¡£¡£¡£¡£¡£
È«¾Ö l3 kv cache pool£ºÔÚ¶àÂֺͳ¬³¤ÉÏÏÂÎÄµÄ agent ³¡¾°Ï£¬£¬£¬£¬£¬ÇëÇó¼äÓµÓм«¸ßµÄ¹²Ïíǰ׺±ÈÀý£¬£¬£¬£¬£¬¿ÉÊǾֲ¿µÄ kv cache ÊÜÈÝÁ¿ÏÞÖÆ£¬£¬£¬£¬£¬ÎÞ·¨µÖ´ïÖª×ãµÄ prefix cache ÖÀÖÐÂÊ£¬£¬£¬£¬£¬ÉõÖÁÔÚ rl batch size ¼«´óµÄÇéÐÎÏ£¬£¬£¬£¬£¬»á±¬·¢´ó×ÚÓÉÓÚÇýÖðµ¼ÖµÄÖØÅÌË㣬£¬£¬£¬£¬Òò´ËÐèÒªÖ§³ÖÈ«¾ÖµÄ l3 kv cache¡£¡£¡£¡£¡£¡£¡£Í¬Ê±£¬£¬£¬£¬£¬forge »¹Í¨¹ý scheduler cost-aware µÄµ÷Àí»úÖÆ£¬£¬£¬£¬£¬È¨ºâÅŶÓÑÓ»ºÂýºÍ´æ´«Êäʱ¼äÀ´¶¯Ì¬Â·ÓÉÇëÇ󣬣¬£¬£¬£¬ÔÚ²»Ê¹ÊµÀý³¬ÔصÄÌõ¼þÏÂ×î´ó»¯»º´æ¾Ö²¿ÐÔ¡£¡£¡£¡£¡£¡£¡£
scalable agent rl Ëã·¨
![]()
ΪÏàʶ¾ö³¬³¤¹ì¼£µÄÐÅÓ÷ÖÅÉÎÊÌⲢȷ±£Îȹ̣¬£¬£¬£¬£¬ÎÒÃÇÉè¼ÆÁËÒ»¸öÓÉÈý²¿·Ö×é³ÉµÄ¸´ºÏ½±Àø£º
1.Àú³Ì½±Àø£¨process reward£©£º¼àÊÓ agent µÄÖÐÐÄÐÐΪ£¨Èç´¦·ÖÓïÑÔ»ìÏý»òÌØ¶¨¹¤¾ßŲÓùýʧ£©£¬£¬£¬£¬£¬Ìṩ÷缯·´Ï죡£¡£¡£¡£¡£¡£¬£¬£¬£¬£¬¶ø²»µ«ÒÀÀµ×îÖÕЧ¹û¡£¡£¡£¡£¡£¡£¡£
2.ʹÃüÍê³Éʱ¼ä½±Àø£º½«Ïà¶ÔÍê³Éʱ¼ä×÷Ϊ½±ÀøÐźš£¡£¡£¡£¡£¡£¡£ÓÉÓÚÕæÊµÑÓ³Ù²»µ«È¡¾öÓÚ token ÌìÉú£¬£¬£¬£¬£¬»¹Êܹ¤¾ßÖ´ÐкÍ×Ó agent ŲÓÃÓ°Ï죬£¬£¬£¬£¬ÕâÄܼ¤Àø agent ×Ô¶¯Ê¹Óò¢ÐÐÕ½ÂÔ¡¢Ñ¡Ôñ×î¶ÌµÄÖ´Ðз¾¶À´¼ÓËÙʹÃü¡£¡£¡£¡£¡£¡£¡£
3.ÓÃÓÚ½µµÍ·½²îµÄºóÐø½±Àø£¨reward-to-go£©£º³¤ÖÜÆÚʹÃüµÄÏ£º±½±ÀøÈÝÒ×Òý·¢¸ßÌݶȷ½²î¡£¡£¡£¡£¡£¡£¡£ÎÒÃÇʹÓà reward-to-go À´±ê×¼»¯»Ø±¨£¬£¬£¬£¬£¬´ó·ùÌá¸ßÁËÐÅÓ÷ÖÅɵľ«¶È£¬£¬£¬£¬£¬ÎȹÌÁËÓÅ»¯Àú³Ì¡£¡£¡£¡£¡£¡£¡£
ѵ³öÒ»¸öÕæÕýºÃÓõÄÄ£×Ó£¬£¬£¬£¬£¬¹¤³Ì¡¢Êý¾Ý¡¢Ë㷨ȱһ²»¿É£¬£¬£¬£¬£¬ÄܸÏÔÚÄêǰ½»³öÕâ·Ý´ð¾í£¬£¬£¬£¬£¬Àë²»¿ª±³ºóÿһλͬÊÂµÄÆð¾¢¡£¡£¡£¡£¡£¡£¡£¿£¿£¿£¿£¿£¿´µ½ÁËÉçÇøºÜÊǶàµÄÕýÏò·´Ïì¸ÐÓ¦·Ç³£¿£¿£¿£¿£¿£¿ªÐÄ£¬£¬£¬£¬£¬×Åʵ m2.5 ÉÐÓкܴóµÄÌáÉý¿Õ¼ä£¬£¬£¬£¬£¬ÄÚ²¿ rl Ò²»¹ÔÚ¼ÌÐøÅÜ£¬£¬£¬£¬£¬ÐÔÄÜÒ²ÔÚÒ»Á¬ÕÇ¡£¡£¡£¡£¡£¡£¡£ÏÖÔÚ£¬£¬£¬£¬£¬m2.5 ÒѾÖÜÈ«¿ªÔ´¡£¡£¡£¡£¡£¡£¡£
hugging face: huggingface.co/minimaxai/minimax-m2.5
github: github.com/minimax-ai/minimax-m2.5
´º½ÚÂíÉϵ½ÁË£¬£¬£¬£¬£¬×£¸÷ÈËÐÂÄê¿ìÀÖ£¡
½Ó´ýת·¢£¬£¬£¬£¬£¬µ«Çë×¢Ã÷À´ÓÉ¡°ÉϺ£¾ÐÅί¡±
ÉϹۺÅ×÷ÕߣºÉϺ£¾ÐÅί
¼ÓÄôópc×îÐÂÕ¹ÍûÆÊÎö±¨¸æ>
