MARLとは? わかりやすく解説

Marl

名前 マール

マルチエージェント強化学習

(MARL から転送)

出典: フリー百科事典『ウィキペディア(Wikipedia)』 (2026/08/15 06:32 UTC 版)

敵対するエージェントの2チームが対決するMARLの実験

マルチエージェント強化学習 (MARL) とは、単一の共有環境下にいる、複数の学習エージェントの挙動を研究する強化学習の一分野である[1]。それぞれのエージェントは、自身の報酬に動機付けられており、自己利益を増大するために行動する。この利益が他のエージェントと相反する環境も存在し、複雑なグループダイナミクスを引き起こす。

マルチエージェント強化学習は、マルチエージェントシステムやゲーム理論と密接に関係しており、特に繰り返しゲームと関係が深い。MARLの研究においては、報酬最大化の理想的アルゴリズムの探求と、より社会学的な概念群を結びつける。単一エージェントの強化学習の研究は一体のエージェントの報酬最大化アルゴリズムに関心を向けるが、MARLでは協調[2]、互恵性[3]、平等性[4]、社会的影響[5]、言語[6]、統計的差別[7]、といった社会学的指標を定量化し、評価する。

定義

単一エージェントの強化学習と同じように、MARLでは、マルコフ決定過程(MDP)の一種としてモデル化することが多い。

  • エージェントの集合:
    混合和設定で、四体のエージェントのそれぞれが異なるゴールに到達しようとしている。それぞれのエージェントの成功はほかのエージェントが道を空けるか否かに依存する。彼らに他のエージェントを支援する直接的な動機はないのであるが[16]。

    複数エージェントが関与する多くの実世界シナリオは協調と競争の両要素を持つ。例えば、複数の自動運転車がそれぞれの経路を計画しているとき、それぞれはバラバラで背反でない利益を持っている。それぞれの車は目的地にたどり着く時間を最小化しようとしているが、全ての車が交通事故を回避したいという共通利害を持っている[17]。

    3人以上のゼロサム設定では、しばしば混合和設定と似たような特性を呈する。それぞれのエージェントのペアは非ゼロの効用和を持つためである。

    混合和設定は囚人のジレンマといった古典的な戦略型ゲーム(行列ゲーム)や、より複雑な逐次社会的ジレンマ(SSDs)、Among Us[18]、Diplomacy[19]、StarCraft II[20][21]といった娯楽用ゲームでも研究できる。

    混合和設定はコミュニケーションと社会的ジレンマを引き起こしうる。

社会的ジレンマ

ゲーム理論では、MARLの多くの研究は、例えば囚人のジレンマ[22]や、チキンゲームや鹿狩りゲームのような、社会的ジレンマを中心に展開している[23]。

ゲーム理論の研究はナッシュ均衡や、単一エージェントがとるべき理想的方策は何かということに関心を向けるが、MARLの研究は、複数エージェントが試行錯誤過程の中でどのように理想的方策を獲得するかということに目を向ける。エージェントを訓練するために使われる強化学習アルゴリズムは自身の報酬の最大化を行うが、エージェントの利益と集団の利益の衝突が活発な研究の主題となっている[24]。

環境のルールを修正する[25]、内発的報酬を追加する[4]など、協調を促す様々なテクニックが研究されてきた。

逐次社会的ジレンマ(SSD)

囚人のジレンマ、チキンゲーム、鹿狩りゲームといった社会的ジレンマは行列ゲームである。それぞれのエージェントは二つの選択肢から一つだけを一度だけ選ぶ、シンプルな2×2の行列で取った行動に応じた報酬が記述できるゲームである。

人間や他の生物では社会的ジレンマはより複雑な傾向がある。エージェントが時間とともに複数の行動をとり、協力と裏切りの区別は行列ゲームほど自明ではない。逐次社会的ジレンマ(SSD)という概念は、その複雑性をモデル化するために、2017年[26] に導入された。SSDの様々な種類を定義し、その中でのエージェントの行動における協調的ふるまいを示す、研究が進行している[27]。

自動カリキュラム

自動カリキュラム[28]とは、特にマルチエージェントの実験において顕著に生じる強化学習の概念である。エージェントが性能を向上させるに従って、エージェントが環境を変化させ、この環境の変化が自身や、他のエージェントに影響を与える。このフィードバックループは、様々な区別される学習の段階を引き起こす。このそれぞれは前の段階に依存している。この学習の積み重なる階層を自動カリキュラムと呼ぶ。自動カリキュラムは特に各エージェントのグループが対立するグループの現在の戦略に対抗しようと競い合う敵対的設定で顕著である,[29]。

かくれんぼゲームは、敵対的設定において自動カリキュラムが生じる分かりやすい例である。この実験では、seekerのチームとhiderのチームが対戦する。一方のチームが新しい戦略を習得するたびに、相手チームはそれに対して最も効果的な対抗策をとるよう自らの戦略を適応させる。hiderが箱を使ってシェルターを作ることを学習すると、seekerはスロープを使ってそのシェルターに侵入することを学習して応答する。それに対し、hiderはスロープをロックすることでseekerに使わせないようにする。するとseekerは、ゲームのグリッチを悪用してシェルター内に侵入する「ボックスサーフィン」によってこれに応答する。学習の各「レベル」は、直前のレベルを前提として現れる創発現象である。これにより、直前の行動に依存した一連の行動の積み重ねが生じる。

強化学習実験における自動カリキュラムは、生命の進化の段階や、人類の文化の発達に喩えられる。進化における主要な段階の一つは20億〜30億年前に起こり、光合成を行う生命体が大量の酸素を生成し始め、大気成分の割合を変化させた[30]。進化の次の段階では、酸素呼吸を行う生命体が進化し、最終的に陸上の哺乳類や人類へとつながった。これらの後半の段階は、光合成の段階によって酸素が広く利用可能になった後に初めて可能となったものである。同様に、人類の文化も、紀元前10000年頃の農業革命によって得られた資源や知見なしには、18世紀の産業革命を迎えることはできなかったとされる[31]。

アルゴリズム

一般に、マルチエージェント強化学習アルゴリズムは、PPOやDQN(深層Q学習)といった一般的な強化学習アルゴリズムを拡張したものである。協調的なシナリオにおいては、これらのアルゴリズムはエージェント間のコミュニケーションを特徴とする。しばしば、エージェントが協調して行動し累積報酬を最大化することを学習できるよう、集中学習・分散実行(Centralized Training, Decentralized Execution, CTDE)という手法が用いられる。

マルチエージェントDQN系

独立Q学習(IQL)

最も単純なアプローチは、他のすべてのエージェントを環境の一部として扱うというもので、独立Q学習(Independent Q Learning)として知られる手法である。しかし、これは非定常性と呼ばれる重大な問題を引き起こす[32]。

非定常性問題

各エージェントは、他のすべてのエージェントを環境の一部として扱う。しかし、各エージェントは自らの方策だけでなく環境そのものも変化させているため、すべてのエージェントにとって一種の「移動ターゲット」現象が生まれ、収束の保証が失われ、環境が非常に混沌としたものになる[32]。

価値分解ネットワーク(VDN)

Dec-POMDPにおける協調的なマルチエージェントDQNに対する最も単純なアプローチは、各エージェントが計算したQ値を合計し、大域的なTD誤差に基づいて最適化を行うというもので、これは価値分解ネットワーク(Value-Decomposition Network)として知られる手法である[33]。具体的には、

  • Stefano V. Albrecht, Filippos Christianos, Lukas Schäfer. Multi-Agent Reinforcement Learning: Foundations and Modern Approaches. MIT Press, 2024. https://www.marl-book.com
  • Kaiqing Zhang, Zhuoran Yang, Tamer Basar. Multi-agent reinforcement learning: A selective overview of theories and algorithms. Studies in Systems, Decision and Control, Handbook on RL and Control, 2021.
  • Yang, Yaodong; Wang, Jun (2020). “An Overview of Multi-Agent Reinforcement Learning from Game Theoretical Perspective”. arXiv:2011.00583 [cs.MA].

出典

  1. ↑ Stefano V. Albrecht, Filippos Christianos, Lukas Schäfer. Multi-Agent Reinforcement Learning: Foundations and Modern Approaches. MIT Press, 2024. https://www.marl-book.com/
  2. ↑ Lowe, Ryan; Wu, Yi (2020). “Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments”. arXiv:1706.02275v4 [cs.LG].
  3. ↑ Baker, Bowen (2020). “Emergent Reciprocity and Team Formation from Randomized Uncertain Social Preferences”. NeurIPS 2020 proceedings. arXiv:2011.05373.
  4. 1 2 Hughes, Edward; Leibo, Joel Z.; et al. (2018). “Inequity aversion improves cooperation in intertemporal social dilemmas”. NeurIPS 2018 proceedings. arXiv:1803.08884.
  5. ↑ Jaques, Natasha; Lazaridou, Angeliki; Hughes, Edward; et al. (2019). “Social Influence as Intrinsic Motivation for Multi-Agent Deep Reinforcement Learning”. Proceedings of the 35th International Conference on Machine Learning. arXiv:1810.08647.
  6. ↑ Lazaridou, Angeliki (2017). “Multi-Agent Cooperation and The Emergence of (Natural) Language”. ICLR 2017. arXiv:1612.07182.
  7. ↑ Duéñez-Guzmán, Edgar; et al. (2021). “Statistical discrimination in learning agents”. arXiv:2110.11404v1 [cs.LG].
  8. ↑ Campbell, Murray; Hoane, A. Joseph Jr.; Hsu, Feng-hsiung (2002). “Deep Blue”. Artificial Intelligence (Elsevier) 134 (1–2): 57–83. doi:10.1016/S0004-3702(01)00129-1. ISSN 0004-3702.
  9. ↑ Carroll, Micah; et al. (2019). “On the Utility of Learning about Humans for Human-AI Coordination”. arXiv:1910.05789 [cs.LG].
  10. ↑ Xie, Annie; Losey, Dylan; Tolsma, Ryan; Finn, Chelsea; Sadigh, Dorsa (November 2020). Learning Latent Representations to Influence Multi-Agent Interaction (PDF). CoRL.
  11. ↑ Clark, Herbert; Wilkes-Gibbs, Deanna (February 1986). “Referring as a collaborative process”. Cognition 22 (1): 1–39. doi:10.1016/0010-0277(86)90010-7. PMID 3709088.
  12. ↑ Boutilier, Craig (17 March 1996). “Planning, learning and coordination in multiagent decision processes”. Proceedings of the 6th Conference on Theoretical Aspects of Rationality and Knowledge: 195–210.
  13. ↑ Stone, Peter; Kaminka, Gal A.; Kraus, Sarit; Rosenschein, Jeffrey S. (July 2010). Ad Hoc Autonomous Agent Teams: Collaboration without Pre-Coordination. AAAI 11.
  14. ↑ Foerster, Jakob N.; Song, H. Francis; Hughes, Edward; Burch, Neil; Dunning, Iain; Whiteson, Shimon; Botvinick, Matthew M; Bowling, Michael H. Bayesian action decoder for deep multi-agent reinforcement learning. ICML 2019. arXiv:1811.01458.
  15. ↑ Shih, Andy; Sawhney, Arjun; Kondic, Jovana; Ermon, Stefano; Sadigh, Dorsa. On the Critical Role of Conventions in Adaptive Human-AI Collaboration. ICLR 2021. arXiv:2104.02871.
  16. ↑ Bettini, Matteo; Kortvelesy, Ryan; Blumenkamp, Jan; Prorok, Amanda (2022). “VMAS: A Vectorized Multi-Agent Simulator for Collective Robot Learning”. The 16th International Symposium on Distributed Autonomous Robotic Systems (Springer). arXiv:2207.03530.
  17. ↑ Shalev-Shwartz, Shai; Shammah, Shaked; Shashua, Amnon (2016). “Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving”. arXiv:1610.03295 [cs.AI].
  18. ↑ Kopparapu, Kavya; Duéñez-Guzmán, Edgar A.; Matyas, Jayd; Vezhnevets, Alexander Sasha; Agapiou, John P.; McKee, Kevin R.; Everett, Richard; Marecki, Janusz; Leibo, Joel Z.; Graepel, Thore (2022). “Hidden Agenda: a Social Deduction Game with Diverse Learned Equilibria”. arXiv:2201.01816 [cs.AI].
  19. ↑ Bakhtin, Anton構文エラー:「etal」を認識できません。 (2022). “Human-level play in the game of Diplomacy by combining language models with strategic reasoning”. Science (Springer) 378 (6624): 1067–1074. Bibcode:2022Sci...378.1067M. doi:10.1126/science.ade9097. PMID 36413172.
  20. ↑ Samvelyan, Mikayel; Rashid, Tabish; de Witt, Christian Schroeder; Farquhar, Gregory; Nardelli, Nantas; Rudner, Tim G. J.; Hung, Chia-Man; Torr, Philip H. S.; Foerster, Jakob; Whiteson, Shimon (2019). “The StarCraft Multi-Agent Challenge”. arXiv:1902.04043 [cs.LG].
  21. ↑ Ellis, Benjamin; Moalla, Skander; Samvelyan, Mikayel; Sun, Mingfei; Mahajan, Anuj; Foerster, Jakob N.; Whiteson, Shimon (2022). “SMACv2: An Improved Benchmark for Cooperative Multi-Agent Reinforcement Learning”. arXiv:2212.07489 [cs.LG].
  22. ↑ Sandholm, Toumas W.; Crites, Robert H. (1996). “Multiagent reinforcement learning in the Iterated Prisoner's Dilemma”. Biosystems 37 (1–2): 147–166. Bibcode:1996BiSys..37..147S. doi:10.1016/0303-2647(95)01551-5. PMID 8924633.
  23. ↑ Peysakhovich, Alexander; Lerer, Adam (2018). “Prosocial Learning Agents Solve Generalized Stag Hunts Better than Selfish Ones”. AAMAS 2018. arXiv:1709.02865.
  24. ↑ Dafoe, Allan; Hughes, Edward; Bachrach, Yoram; et al. (2020). “Open Problems in Cooperative AI”. NeurIPS 2020. arXiv:2012.08630.
  25. ↑ Köster, Raphael; Hadfield-Menell, Dylan; Hadfield, Gillian K.; Leibo, Joel Z. “Silly rules improve the capacity of agents to learn stable enforcement and compliance behaviors”. AAMAS 2020. arXiv:2001.09318.
  26. ↑ Leibo, Joel Z.; Zambaldi, Vinicius; Lanctot, Marc; Marecki, Janusz; Graepel, Thore (2017). “Multi-agent Reinforcement Learning in Sequential Social Dilemmas”. AAMAS 2017. arXiv:1702.03037.
  27. ↑ Badjatiya, Pinkesh; Sarkar, Mausoom (2020). “Inducing Cooperative behaviour in Sequential-Social dilemmas through Multi-Agent Reinforcement Learning using Status-Quo Loss”. arXiv:2001.05458 [cs.AI].
  28. ↑ Leibo, Joel Z.; Hughes, Edward; et al. (2019). “Autocurricula and the Emergence of Innovation from Social Interaction: A Manifesto for Multi-Agent Intelligence Research”. arXiv:1903.00742v2 [cs.AI].
  29. ↑ Baker, Bowen; et al. (2020). “Emergent Tool Use From Multi-Agent Autocurricula”. ICLR 2020. arXiv:1909.07528.
  30. ↑ Kasting, James F; Siefert, Janet L (2002). “Life and the evolution of earth's atmosphere”. Science 296 (5570): 1066–1068. Bibcode:2002Sci...296.1066K. doi:10.1126/science.1071184. PMID 12004117.
  31. ↑ Clark, Gregory (2008). A farewell to alms: a brief economic history of the world. Princeton University Press. ISBN 978-0-691-14128-2
  32. 1 2 Tan, Ming (1993). “Multi Agent Reinforcement Learning: Independent vs Cooperative Agents”. International Conference on Machine Learning.
  33. ↑ Sunehag, Peter; Lever, Guy; Gruslys, Audrunas; Czarnecki, Wojciech Marian; Zambaldi, Vinicius; Jaderberg, Max; Lanctot, Marc; Sonnerat, Nicolas et al. (2018-07-09). “Value-Decomposition Networks For Cooperative Multi-Agent Learning Based On Team Reward”. Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems. AAMAS '18 (Richland, SC: International Foundation for Autonomous Agents and Multiagent Systems): 2085–2087.
  34. 1 2 3 4 5 6 7 8 Li, Tianxu; Zhu, Kun; Luong, Nguyen Cong; Niyato, Dusit; Wu, Qihui; Zhang, Yang; Chen, Bing (2021). “Applications of Multi-Agent Reinforcement Learning in Future Internet: A Comprehensive Survey”. arXiv:2110.13484 [cs.AI].
  35. ↑ Le, Ngan; Rathour, Vidhiwar Singh; Yamazaki, Kashu; Luu, Khoa; Savvides, Marios (2021). “Deep Reinforcement Learning in Computer Vision: A Comprehensive Survey”. arXiv:2108.11510 [cs.CV].
  36. ↑ Moulin-Frier, Clément; Oudeyer, Pierre-Yves (2020). “Multi-Agent Reinforcement Learning as a Computational Tool for Language Evolution Research: Historical Context and Future Challenges”. arXiv:2002.08878 [cs.MA].
  37. ↑ Killian, Jackson; Xu, Lily; Biswas, Arpita; Verma, Shresth; et al. (2023). Robust Planning over Restless Groups: Engagement Interventions for a Large-Scale Maternal Telehealth Program. AAAI.
  38. ↑ Krishnan, Srivatsan; Jaques, Natasha; Omidshafiei, Shayegan; Zhang, Dan; Gur, Izzeddin; Reddi, Vijay Janapa; Faust, Aleksandra (2022). “Multi-Agent Reinforcement Learning for Microprocessor Design Space Exploration”. arXiv:2211.16385 [cs.AR].
  39. ↑ Li, Yuanzheng; He, Shangyang; Li, Yang; Shi, Yang; Zeng, Zhigang (2023). “Federated Multiagent Deep Reinforcement Learning Approach via Physics-Informed Reward for Multimicrogrid Energy Management”. IEEE Transactions on Neural Networks and Learning Systems PP (5): 5902–5914. arXiv:2301.00641. doi:10.1109/TNNLS.2022.3232630. PMID 37018258.
  40. ↑ Ci, Hai; Liu, Mickel; Pan, Xuehai; Zhong, Fangwei; Wang, Yizhou (2023). Proactive Multi-Camera Collaboration for 3D Human Pose Estimation. International Conference on Learning Representations.
  41. ↑ Vinitsky, Eugene; Kreidieh, Aboudy; Le Flem, Luc; Kheterpal, Nishant; Jang, Kathy; Wu, Fangyu; Liaw, Richard; Liang, Eric; Bayen, Alexandre M. (2018). Benchmarks for reinforcement learning in mixed-autonomy traffic (PDF). Conference on Robot Learning.
  42. ↑ Tuyls, Karl; Omidshafiei, Shayegan; Muller, Paul; Wang, Zhe; Connor, Jerome; Hennes, Daniel; Graham, Ian; Spearman, William; Waskett, Tim; Steele, Dafydd; Luc, Pauline; Recasens, Adria; Galashov, Alexandre; Thornton, Gregory; Elie, Romuald; Sprechmann, Pablo; Moreno, Pol; Cao, Kris; Garnelo, Marta; Dutta, Praneet; Valko, Michal; Heess, Nicolas; Bridgland, Alex; Perolat, Julien; De Vylder, Bart; Eslami, Ali; Rowland, Mark; Jaegle, Andrew; Munos, Remi; Back, Trevor; Ahamed, Razia; Bouton, Simon; Beauguerlange, Nathalie; Broshear, Jackson; Graepel, Thore; Hassabis, Demis (2020). “Game Plan: What AI can do for Football, and What Football can do for AI”. arXiv:2011.09192 [cs.AI].
  43. ↑ Chu, Tianshu; Wang, Jie; Codec├á, Lara; Li, Zhaojian (2019). “Multi-Agent Deep Reinforcement Learning for Large-scale Traffic Signal Control”. IEEE Transactions on Intelligent Transportation Systems 21 (3): 1086. arXiv:1903.04527. Bibcode:2020ITITr..21.1086C. doi:10.1109/TITS.2019.2901791.
  44. ↑ Belletti, Francois; Haziza, Daniel; Gomes, Gabriel; Bayen, Alexandre M. (2017). “Expert Level control of Ramp Metering based on Multi-task Deep Reinforcement Learning”. arXiv:1701.08832 [cs.AI].
  45. ↑ Ding, Yahao; Yang, Zhaohui; Pham, Quoc-Viet; Zhang, Zhaoyang; Shikh-Bahaei, Mohammad (2023). “Distributed Machine Learning for UAV Swarms: Computing, Sensing, and Semantics”. arXiv:2301.00912 [cs.LG].
  46. ↑ Xu, Lily; Perrault, Andrew; Fang, Fei; Chen, Haipeng; Tambe, Milind (2021). “Robust Reinforcement Learning Under Minimax Regret for Green Security”. arXiv:2106.08413 [cs.LG].
  47. ↑ Leike, Jan; Martic, Miljan; Krakovna, Victoria; Ortega, Pedro A.; Everitt, Tom; Lefrancq, Andrew; Orseau, Laurent; Legg, Shane (2017). “AI Safety Gridworlds”. arXiv:1711.09883 [cs.AI].
  48. ↑ Hadfield-Menell, Dylan; Dragan, Anca; Abbeel, Pieter; Russell, Stuart (2016). “The Off-Switch Game”. arXiv:1611.08219 [cs.AI].
  49. ↑ Hernandez-Leal, Pablo; Kartal, Bilal; Taylor, Matthew E. (2019-11-01). “A survey and critique of multiagent deep reinforcement learning” (英語). Autonomous Agents and Multi-Agent Systems 33 (6): 750–797. arXiv:1810.05587. doi:10.1007/s10458-019-09421-1. ISSN 1573-7454.

泥灰土

(MARL から転送)

出典: フリー百科事典『ウィキペディア(Wikipedia)』 (2020/12/12 00:36 UTC 版)

ナビゲーションに移動 検索に移動
フランスのノルマンディーの泥灰土
石灰と粘土の比率と分類

泥灰土(でいかいど、marl)とは、粘土質物質と石灰もしくは炭酸カルシウム(方解石など)の混合物で、非常にまれにドロマイトを含む堆積物。マールとも言う。

概要

炭酸塩鉱物を35-65%含み、残りが粘土からなる堆積物を泥灰土としており、これが固結したものが泥灰岩(marlstone, marlite)と定義されている[1]。粘土分の比率が増加すると石灰質粘土(calcareous clay)となり、減少すると粘土質石灰岩(argillaceous limestone)となる。

用途

土壌改良
マールは1800年代にニュージャージー州中部で土壌改良用途に広く採掘された。酸性土壌の中和、保水性の改善目的で畑に使用される。
セメント
ポルトランドセメントの原料とされる。そのままの組成で使える泥灰岩は、セメント岩(セメントロック)と呼ばれる。一般に酸化鉄やマグネシウムが多すぎるなど成分にばらつきがあるという欠点がある[2]。

出典

  1. ^ Pettijohn, F. J. (1957). Sedimentary Rocks (2nd ed.). New York: Harper & Brothers. OCLC 551748. p.410
  2. ^ 岡崎清, 新居善三郎, 素木洋一, 吉木文平, 稲生謙次, 井本文夫, 内藤雅夫, 近藤連一、「セラミック原料解説集 チ」『窯業協會誌』 1965年 73巻 835号 p.C212-C218, doi:10.2109/jcersj1950.73.835_C212, 日本セラミックス協会


英和和英テキスト翻訳

英語⇒日本語日本語⇒英語

辞書ショートカット

すべての辞書の索引

「MARL」の関連用語

MARLのお隣キーワード
検索ランキング

   

英語⇒日本語
日本語⇒英語
   



MARLのページの著作権

   
日外アソシエーツ株式会社日外アソシエーツ株式会社
Copyright (C) 1994- Nichigai Associates, Inc., All rights reserved.
ウィキペディアウィキペディア
All text is available under the terms of the GNU Free Documentation License.
この記事は、ウィキペディアのマルチエージェント強化学習 (改訂履歴)、泥灰土 (改訂履歴)の記事を複製、再配布したものにあたり、GNU Free Documentation Licenseというライセンスの下で提供されています。 Weblio辞書に掲載されているウィキペディアの記事も、全てGNU Free Documentation Licenseの元に提供されております。

©2026 GRAS Group, Inc.RSS