APP下载

A single-task and multi-decision evolutionary game model based on multi-agent reinforcement learning

2021-07-26MAYeCHANGTianqingandFANWenhui

MA Ye, CHANG Tianqing, and FAN Wenhui

1. Academy of Army Armored Force, Beijing 100072, China;

2. Department of Automation, Tsinghua University, Beijing 100084, China

Abstract: In the evolutionary game of the same task for groups,the changes in game rules, personal interests, the crowd size,and external supervision cause uncertain effects on individual decision-making and game results. In the Markov decision framework, a single-task multi-decision evolutionary game model based on multi-agent reinforcement learning is proposed to explore the evolutionary rules in the process of a game. The model can improve the result of a evolutionary game and facilitate the completion of the task. First, based on the multi-agent theory, to solve the existing problems in the original model, a negative feedback tax penalty mechanism is proposed to guide the strategy selection of individuals in the group. In addition, in order to evaluate the evolutionary game results of the group in the model, a calculation method of the group intelligence level is defined. Secondly, the Q-learning algorithm is used to improve the guiding effect of the negative feedback tax penalty mechanism. In the model, the selection strategy of the Q-learning algorithm is improved and a bounded rationality evolutionary game strategy is proposed based on the rule of evolutionary games and the consideration of the bounded rationality of individuals. Finally, simulation results show that the proposed model can effectively guide individuals to choose cooperation strategies which are beneficial to task completion and stability under different negative feedback factor values and different group sizes, so as to improve the group intelligence level.

Keywords: multi-agent, reinforcement learning, evolutionary game, Q-learning.

1. Introduction

Reinforcement learning is a model of machine learning. It can actively sense the environment via different behaviors or actions, evaluate the actions and adjust subsequent actions. It is a learning technology mapping different environmental states into actions [1]. Reinforcement learning mainly aims to choose the optimal action of the agents when they complete goals. It is widely used in robot control systems [2,3], intelligent decision-making[4,5], nonlinear optimal control [6-8] and other fields. A single agent generally has no proper decision-making ability or the ability to sense the environment, so it cannot respond to complex actual problems. Therefore, the concept of multi-agent was proposed at the end of the 20th century. It is a collection of multiple agents and also called multi-agent system. It belongs to the forefront of distributed artificial intelligence and is mainly used to explore the coordination, cooperation, communication, and conflicts among agent groups [9].

Reinforcement learning is also applied in multi-agent learning. In recent years, the combination of multi-agent learning and reinforcement learning has become a research hotspot [10-12]. One of the prominent achievements is AlphaGO, the Chinese game of go system based on reinforcement learning, which has defeated the top human players in the game and shows a great advantage. It immediately attracts the attention of all walks of life and more researchers have participated in the field of multiagent reinforcement learning [13].

Game is a theory of action, used to study the strategic choices driven by multiple interests among multiple individuals [14]. Due to the development of the game theory,many research methods have been combined with the game theory, such as genetic algorithms [15], particle swarm optimization [16], and multi-agent systems. The idea of cooperation among agents in a multi-agent system can well reflect the game process of individuals in social groups in their work. The multi-agent system has been gradually introduced into the game field and achieved better results [17,18]. The game involves not only the action of an agent, but also the states of other agents, which increase the complexity of the system and learning, so the convergence rate cannot be guaranteed. The introduction of reinforcement learning can better guide the process of the game. Bendor et al. [19] used reinforcement learning to study the steady-state convergence problem of games.Jacob et al. [20] improved the reinforcement learning algorithm and solved the poor player strategy selection problem in two-player tasks. As one of the reinforcement learning algorithms, the Q-learning algorithm does not require the dynamic environment and the characteristics suitable for long-scenario tasks in advance, so it is widely used in the game field. For example, Littman et al. [4]proposed the minimax algorithm. This algorithm can well solve the two-player zero-sum game, but it cannot be applied in games with more than two players. Jun et al. [21]proposed a random game Q-learning algorithm to search for the optimal strategy. In addition, according to the different effects of the reward function in the game task, the algorithm can be divided into three different types: fully cooperative, fully competitive, and mixed types [22]. The reward functions for different agents in the fully cooperative algorithm are the same and can be used in multi-intelligence systems with the same goal. The classic algorithms are distributed Q-learning algorithms [23] and team Q-learning algorithms [24]. The agents in the fully competitive algorithm are in a state of competition with each other in order to maximize their own returns while minimizing others’ returns. Its classic algorithms include the Minimax-Q learning algorithm [25] and the Nash-Q learning algorithm [26]. The return function in the hybrid algorithm is not related to each other and there is no deterministic rule. It is suitable for the study on the equilibrium solution in the game theory. Its classic algorithms include the fuzzy Q-learning algorithm [27] and the correlated Q-learning algorithm [28].

In the study, through the full consideration of the advantages of multi-agents and reinforcement learning in the field of gaming, a single-task multi-decision evolutionary game model is proposed based on multi-agent reinforcement learning to explore the evolution of groups.This model combines the evolutionary game theory to perform a three-player evolutionary game. A negative feedback tax penalty mechanism is proposed and an improved Q-learning algorithm is used to optimize the effect of this mechanism. The algorithm in this paper can take into account the individual’s incomplete rationality.In addition, a calculation method of group intelligence level is defined to evaluate the results of the group evolutionary game. The simulation results show that the introduction of the reinforcement learning algorithm can improve the guiding role of the negative feedback tax penalty mechanism and promote the evolution of decisionmaking group towards the direction of cooperation.

2. Key technical principles of the model

This paper proposes an evolutionary game model based on multi-agent reinforcement learning to explore the evolution rules in the game process. The key technical principles involved in the model are described below.

2.1 Evolutionary game theory

The evolutionary game theory is an extension of the classical game theory. The significant difference between them is that individuals in the evolutionary game theory can be non-rational [29,30]. The main research object of evolutionary games is the group game, namely, the game of multiple individuals. In the game, the interactive characteristics of the group strategy are described below. Firstly,there is no identity difference among individuals and the only difference among individuals is the selected strategy.Secondly, the group has only one strategy set and the number of strategies is limited. Individuals choose strategies from the strategy set. Individuals adopting the same strategy have the same rewards and their rewards depend entirely on the currently selected strategy. Thirdly, the rewards of each strategy are related to the number or proportion of the choice of the corresponding strategy.

The evolutionary game model consists of two main parts: group game and group state update, as shown in Fig. 1. The group game consists of three parts: the number of individuals, the set of strategies, and the reward utility function. Let the total number of individuals in the group beN, and the strategy set bes={1,2,···,m} (mis the total number of strategies);xiis the total number of individuals who choose strategyi∈S. The group state isThe group’s state set is The reward of each individual is represented by the reward utility functionUi∈X→R, which corresponds to its chosen strategy. It represents the mapping from the state to the set of real numbers.

The change of the group state caused by the change of group time is the core of the evolutionary game. The group’s decision-making action in the game process can be analyzed according to the change of the group state[31].

2.2 Framework of game learning

In the evolutionary game, the group state is updated by adjusting its own strategy according to the game learning rules, which generally include individual game rules and information comparison rules with other individual strategies and rewards. The above process is also called game learning [32]. In each time stept, each individualvi∈ν will continuously update its own strategysi(t)∈Siduring the game cycle, wheres(t)=(s1(t),s2(t),···,sm(t))∈Sands(t) is the current strategies combination of all individuals. Each individual gets rewards πi(t)=Ui(s(t)).Therefore, the discrete-time evolutionary game is defined as a triple Γ=(ν,{Si|vi∈ν},{Ui|vi∈ν}). The framework of the game learning is shown in Fig. 2.

Fig. 2 Framework of game learning

The game learning rules are generally expressed as

According to (1), all information utilized by each individualhas its own historical strategiesother individuals’ historical strategiesits own reward functionUiand its own historical rewardsAmong them, the individual learning ruleHican be divided into a deterministic function or a random function according to the situation. Each individual can determine the next strategy according to rulessi(t+1)∈Si. The above learning rules are based on the assumption that all individuals are completely rational and can obtain all game information. In reality, individuals may not be consistent with the above assumptions. Therefore, the learning rules for bounded rationality and limited information acquisition ability are expressed as

In other words, the memory ability of each individual is changed from infinite memory ability to limited memory ability, which is closer to the reality.

2.3 Learning framework of evolutionary game based on multi-agent

Multi-agent is introduced into the evolutionary game and each individual in the group is treated as an agent. Each agent can interact with each other and freely choose a strategy. Their own models, methods, and knowledge bases form the basic agent structure. When an agent is learning a game, it is defined as the main decision agent.Through the agent structure, each agent can extract the information required for its strategy selection from the environment and other individuals during game learning and simultaneously store the information into the knowledge base to establish its model library. Finally, based on the information in the method library and game learning information, a comprehensive analysis is performed to complete the strategy selection. The remaining agents are ordinary agents responsible for providing information to the main decision agent. The framework of the evolutionary game based on multi-agent is shown in Fig. 3.

Fig. 3 Framework of evolutionary game based on multi-agent

2.4 Evolutionary game flow based on multi-agent and reinforcement learning

Reinforcement learning is a special branch of machine learning and shares the features of supervised learning and unsupervised learning. Its core is the information interaction between agents and the environment [33,34].The mathematical framework of reinforcement learning is based on the Markov decision process (MDP). The Markov process consists of five key parts:

(ii) The strategy adopted by the agent for state transfer iss(t)=(s1(t),s2(t),···,sm(t))∈S.

(iii) The agent’s transition probability from the statexto the statex′according to the strategysis

(iv) The probability that the agent who transits from the statexto the statex′according to the strategyscan obtain the reward is

(v) The discount factor controlling the reward is γ.

In the process of reinforcement learning, the agent will get corresponding rewards or rewards after the game learning is completed. In general, if the strategy is good,the reward is positive, otherwise it is negative. The agent always hopes to get the maximum reward. The calculation formula of the reward is as follows:

whererTis there ward for the agent’stransi tion from one state to anotherstate within time stepT.If thetaskperformed is a continuous task without a final state, a discount factor γ needs to be introduced to maximize the reward. The discount factor ranges from 0 to 1. Then, the calculation formula of the rewards can be expressed as

The purpose of reinforcement learning is to find the optimal strategy that enables each state of the agent to achieve the correct action. A value function is needed to represent the optimal degree of the agent in a specific state under the strategy π. The hypothetical value function is denoted asV(s), which is the state value under a certain strategy. The function is expressed as

Equation (5) represents the expectation of reward under the strategy π and statex. Substituting (4) into (5)gives

In order to represent the optimal degree of a particular action selected by an agent in a particular state under the strategy π, its state-action value function is defined as theQfunction as follows:

Equation (7) represents the expectation of reward for the actionaunder the strategy π and statex. Substituting (4) into (7) gives

In the evolutionary game process based on multi-agent reinforcement learning, when the agent chooses a strategy, if the environment gives the positive feedback (the reward value is good), the probability that the agent chooses the same strategy will increase in the next round,otherwise it will decrease. Therefore, the decision-making agent will acquire knowledge, learn from the acquired knowledge and the feedback given by the environment, and select a strategy. The flow of the evolutionary game based on multi-agent reinforcement learning is shown in Fig. 4.

Fig. 4 Evolutionary game flow based on multi-agent reinforcement learning

3. Original model

In order to better describe the cooperative relationship and evolution rules in a group, a calculation method of the group intelligence level is defined. Firstly, it is assumed that each agent in the group has an intelligence levelI,I∈[0,1]. The intelligence level of the group isCI,which is the sum of the intelligence levels of individual agents in the group,When the individual agents in the group are playing an evolutionary game, the intelligence level of the group will be changed. Then, the intelligence level of the group iswhereΔIis the intelligent variation generated when individual agents participating in evolutionary games choose different strategies. According to this evolutionary game model, it is assumed that any three individual agents, namely,Agenti, Agentjand Agentk, respectively, have intelligence levels ofIi,Ij, andIk. When they choose different strategies, intelligent variations are generated. The calculation formulas of intelligent variation for cooperation,competition and inaction strategies are provided.

The formula for calculating the intelligent variation of the cooperation strategy is

The formula for calculating the intelligent variation of the competition strategy is

The formula for calculating the intelligent variation of the inaction strategy is

The intelligence level per capita is defined as the ratio of the group intelligence level to the total number of peopleN. According to (9)-(11), in the evolutionary game, the group intelligence levelCImay be any value higher or lower than the sum of individual agent intelligence levelsAccording to the model and the definition of the intelligence level, when all the individual agents participating in the task adopt a cooperation strategy, the intelligence level of the group is the highest. The group has the lowest level of intelligence when the inaction strategy is adopted. When a competition strategy is adopted, the group has the medium intelligence level. The significance of defining the level of group intelligence is to provide an evaluation method for different results produced by different strategies adopted by the group in the process of game evolution. The intelligence level of the group is used as an indicator to measure the overall performance of the decision-making group when completing a task. Based on changes of the group intelligence level, the evolutionary rule of the game and the efficiency of collaboration between groups can be quantitatively analyzed. The changes of the group intelligence level imply that a decision-making individual not only applies the strategy in the threeperson game, but also brings it into work and interaction with other individuals. In other words, an individual strategy is consistent with the behavior of an individual.

In order to explore the evolution rule of individual decision in a group, a simple original evolutionary game model is constructed. The original model is designed for an extreme situation and can more intuitively reflect the influences of key parameters in the model on the selection of the individual strategy and the evolution of the group intelligence level. The model assumes that a group completes a task together. The total number of individuals in the group isNand each individual is regarded as an agent. The cost of completing the task isCand the reward isR, whereR>C. Each individual agent in the group can freely choose its own strategy during the evolutionary game and obtain the right to participate in the task by playing the game through the selected strategy. The set of strategies can be divided into three types: cooperation strategies (those who adopt the strategies are called cooperator, referred to asCo), competition strategies (those who adopt the strategies are called defender, referred to asD), and inaction strategies (those who adopt the strategies are called loner, referred to asL). After the individual agents determine their own strategies, the group will sequentially perform a non-repeating three-person random game and the winners in the game will jointly complete the task. The judgment results of game rules are shown in Table 1. The combination order is not considered in strategy combinations. For example, the combinationsCo Co D,Co D Co, andD Co Couse the same game rules. The game rule is extended from the two-person game decision rule (when the strategy combination isCoCo, both win; when the strategy combination isCoD,Dwins; whenCoL,Cowins, when the strategy combination isDD, a random one wins; when the strategy combination isDL,Dwins; when the strategy combination isLL, no one wins) and stipulates that two of the three are randomly selected to play the two-person game. The winner and the remaining person continue to play the game to select the final three-person game winner. The individual agent that chooses a competitive strategy can only win once, and the two of the three who choose the same strategy have the priority to play the game. For example,when a three-person game is played, two persons choose a cooperative strategy and one person chooses a competitive strategy. The two persons who choose a cooperative strategy will play the game first. Both of them will win,and then play a game with the remaining agent who chooses a competitive strategy. The agent will arbitrarily choose because we only care the number of the final winners. Whoever wins does not affect the number of the final winners. In the end, the winners of the three-person game are one of the agents that chooses the cooperative strategy and the one that chooses the competitive strategy. In summary, this rule ensures that the final result of the strategy combination is not affected by the sequence of the combination. For example, the winners of the combinations ofCoCoD,CoDCo, andDCoCoare allCo D(any one ofCo).

Table 1 Game rules decision tables

In the model, it is assumed that the cost of completing the task is borne by all the agents participating in the task,whereas the rewards of the task are shared equally by all individual agents in the group. In other words, a “freerider” action that an individual agent who does not participate in the task shares the reward is allowed. If the number of individual agent participating in the task isa,at timet, the rewards that can be obtained by the individual agent participating in the task are expressed as

The rewards obtained by an individual agent who does not participate in the task can be expressed as

The evolutionary game learning process is described as follows. After one round of the game is completed, each agent will randomly select a certain number of subgroups for strategic comparison learning. The individual learning agent is also the main decision-making agent. If its own reward is less than the minimum value of the subgroup, the individual agent will copy the strategy of the individual agent with the largest reward in the subgroup.When all individual agents have completed the game learning, they will start the next round of the game and continue to advance the evolutionary game process until the game ends.

According to the model settings, the key parameters that affect the reward and game learning process include task costCand task rewardR. In order to facilitate the description of the relationship between rewards and cost,a simulation experiment is performed on the evolutionary game model with the reward-cost ratioR/C. In real life, the individual’s ability to interact with other individuals is limited, so the number of subgroups is set to 4.The total number of individual agentsNin the group is set to be 500 and the evolutionary game involves 100 rounds. The experiment is repeated 300 times. In the evolutionary game process, the proportions of individual agents choosing different strategies in the group and the group intelligence level under different reward-cost ratios are shown in Fig. 5 (for a clearer display of the changes in different reward-cost ratios, enlarged parts are added).

Fig. 5 Proportions of different strategies and the intelligence level of the group in the evolutionary game

Fig. 5(a)-Fig. 5(c) show the changing trend of the proportions of individuals choosing different strategies in groups with different reward-cost ratios. In the evolutionary game, under different reward-cost ratios, the proportion of choosing cooperation and competition strategies shows an obvious downward trend, whereas the proportion of choosing the inaction strategy shows an obvious upward trend. For the cooperation strategy, when the reward-cost ratios are at a small or intermediate value(R/C=1.5,R/C=2 orR/C=2.5), the proportion of the cooperation strategy in the population is suppressed. When the reward-cost ratioR/Cis large (R/C=3), the proportion of the cooperation strategy in the population has an advantage. The smaller the selected proportion of the cooperation strategy in the group is, the larger the selected proportions of the inaction strategy and the competitive strategy in the group are. When the evolution is stable,the intelligence level of the group with different rewardcost ratios shows a significant decline (see Fig. 5(d)).When the reward-cost ratio is 3, the intelligence level of the group is the highest and the group collaboration effect is also the best.

In summary, due to the acquiescence of the “free-rider”action, the task income is unfairly distributed and the individual agents in the group are always prone to choose the inaction strategy regardless of the change in the reward-cost ratio, so as to ensure their own rewards. However, the task participation rate remains to be at a low level. Due to the sharp decline in the proportion of choosing cooperation strategies, the group intelligence level will be greatly reduced, thus negatively affecting the completion of tasks. In the setting of the original model,when all individuals choose a cooperation strategy, the group has the highest intelligence level and the best task completion effect, and the “free-rider” action does not promote the group to evolve towards the cooperation direction. Therefore, the model needs to be improved in such a way that it can guide the evolutionary game direction of the group to the ideal situation and reduce the proportion of choosing the inaction strategy.

4. Improved model based on the negative feedback tax penalty mechanism

The simulation results of the original model indicate that it is necessary to restrict the “free-rider” action in the group. Therefore, we optimize the reward rules of individual agents, increase taxes to appropriately reduce the rewards of individuals who do not participate in the task,and reward the taxes to the individual agents who participate in the task. In this way, the group is guided to evolve towards the cooperation direction and the group intelligence level is improved. In reality, the formulator of the reward rules may be the leaders of enterprises, institutions and government departments. According to (12)and (13), individuals who have not participated in the task can obtain rewards without any cost. In order to punish the individual agents who do not participate in the task, the model increases the tax rateT(0≤T≤1) to tax the individuals who do not participate in the task and transfer the tax equally to the individuals who participate in the task. In this way, the secondary distribution of the rewards is realized. The smaller the value ofTis, the lighter the punishment effect on the “free-riding” action is. Therefore, the rewards of individuals who are not involved in the task are higher than the rewards of individuals who participate in the task and the number of agents participating in the task decreases. This process is a negative feedback process. The purpose of introducing a negative feedback tax rateTis to reduce the “free-riding”action through the evolutionary game learning process from others and increase the task participation rate. Then,the group evolves towards the cooperative direction. If the punishment is too large, it brings out higher costs and is not conducive to the evolution stability of the group.Therefore, the value of the tax rateTdepends on the result of the evolutionary game. If the number of individuals participating in the task isa, at timet, the rewards of the individuals participating in the task are expressed as

The rewards of individuals who have not participated in the task are expressed as

The model proposes a tax rate penalty mechanism with negative feedback characteristics as follows:

whereNis the total number of people in the group;ais the number of people participating in the task and related to the result of each round of evolutionary games; β is a negative feedback factor, a preset parameter of external forces (its value is about 1 based on the consideration of the actual tax rate). Fig. 6 shows the variation of the tax rateTwith the number of participants underN= 30 and β=1, 0.5 and 1.1.

Fig. 6 Variation of the tax rate T with the number of participants

As shown in Fig. 6, when β=1, the tax rateTis the standard negative feedback and its value decreases as the participation rate increases. When β<1, the penalty effect of the tax rateTis gradually weakened. With the decrease in the number of people participating in the task decreases, the value ofTdecreases and the decreasing rate ofTalso decreases. When β>1, the penalty effect of the tax rateTis gradually increased. The value ofTdecreases significantly with the increase in the number of participants, and the decreasing rate ofTgradually increases. Whentis 0, the model under the negative feedback tax penalty mechanism becomes the original model and the original model can be regarded as a special case.

The simulation is performed with the improved model under the negative feedback tax penalty mechanism. The value of β is about 1, so the six values of β are respectively set as: 0.96, 0.98, 1, 1.02, 1.04, and 1.06. According to the analysis results of the original model, the reward-cost ratio of 3 is more conducive to the guidance of the cooperation strategy. Therefore, the reward-cost ratio is set as 3. The number of individual agentsNin the group is 500. The evolutionary game involves 100 rounds. The evolution of the proportions of individual agents choosing different strategies in the groups under different β values and the group intelligence level are shown in Fig. 7.

Fig. 7 Proportions of different strategies and the intelligence level of the group in the evolutionary game

Fig. 7(a)-Fig. 7(f) show the changing trend of the proportions of individual agents choosing various strategies in groups under different reward-cost ratios. When the negative feedback factor β=1.06, the proportion of individual agents choosing a cooperation strategy is increasing significantly and eventually stabilized at a higher level. When the negative feedback factor β is 0.96, 0.98,1, 1.02, and 1.04, the proportion of individual agents choosing a cooperation strategy firstly shows a temporary upward trend, then significantly declines and is finally remained at a lower level. When β=1.06, the best guidance effect on the cooperation strategy is realized and the proportion of individual agents choosing a cooperation strategy is greatly increased. When other β values are set, during the learning process, the individual agent gradually finds that the rewards of inaction strategies in the task are more advantageous than the rewards of other strategies, and turns to other strategies that are more beneficial to the rewards. When the evolution is stable, cooperation strategies are rarely used. The changing trend of the proportion of the competitive strategy is opposite to that of the cooperation strategy. Whenβ=1.06, the competitive strategy is almost not adopted.When other β values are taken, the competitive strategy is the main strategy adopted by individual agents. The inaction strategy has the least probability to be adopted under the new game learning rules.

As shown in Fig. 7(g) and Fig. 7(h), when the evolution is stable, under different β values except β=1.06,the group intelligence level firstly increases significantly and then remains stable. The group intelligence level under other β values decreases significantly. The change of the intelligence level is related to the trend of the proportions of cooperation strategies. If the proportion of individual agents choosing a competition strategy shows an upward trend, the group intelligence level also shows an upward trend.

In a word, the game learning rules of the negative feedback tax penalty mechanism play a more significant role in limiting the adoption of inaction strategies. Regardless of the value of β, the proportion of individual agents choosing an inaction strategy is almost 0, because the secondary distribution of rewards is not good for agents who are not involved in the task. Under the majority ofβ values, the dominant strategy is the competition strategy.Compared with the original model, the improved model increases the number of people participating in the task,but it has a certain inhibitory effect on the intelligence level of the group. When β=1.06, the dominant strategy is the cooperation strategy. The effect of guiding the group to evolve towards the cooperation mode is the best and the group intelligence level is also improved. When the value of β is less than 1.06, the penalty effect is weakened. Although the group has a tendency to choose a cooperation strategy in the initial stage of the evolution, it is eventually replaced by a competition strategy, which limits the number of people who actually participate in the task and is not conducive to the improvement of the group intelligence level. It can be seen that although the negative feedback tax punishment mechanism increases the number of people participating in the task and has a certain chance to change the group evolution towards the cooperation direction, the dominant strategy is still the competition strategy. In other word, the negative feedback tax punishment mechanism cannot always guide the evolutionary game direction of the group towards the ideal situation. In reality, the formulators of the reward rule need to reasonably formulate the relevant parameters in the negative feedback tax penalty mechanism in order to better guide the group.

5. Improved model based on reinforcement learning algorithms

In the evolutionary game of individual agents in the group, an agent continuously exchanges information with other agents and makes decisions. The agent has a certain ability of autonomous learning. If the model is improved by combining the evolutionary game process with reinforcement learning, it can better guide the direction of the evolutionary game and realize a more ideal evolutionary situation. In the original model, all individual agents are assumed to be fully rational individuals. In reality,game individuals may not always be completely rational.Players sometimes do not follow the rules of game learning when choosing strategies. Therefore, in the improved model based on reinforcement learning, the bounded rationality of individuals will be reflected to some degree.

The original model is built in the Markov decision framework and recorded as the Markov process in a discrete finite state, <S,A,r,p>, whereSandAare respectively discrete state space and action space,ris the reward function of the agent individual. When the individual participates in the task,ris determined by (14).When the individual is not involved in the task,ris determined by (15).pis a transition function determining the transition from one state to another when an individual chooses a certain strategy. However, in the evolutionary game process, the transfer function of the model is unknown, so a special reinforcement learning method,Q-learning algorithm, is required. It can learn without the known transfer function and be suitable for the combination with evolutionary games. In the Q-learning algorithm, the state value is not considered, but the value of the state-action pairQ(s,a), namely, the role of selecting the actionain a certain states, should be considered. TheQvalue is updated from time 1. At timet, theQvalue of timet-1 is updated according to the following formula:

where α∈[0,1] is the learning rate; γ is the discount factor;standatare the states and behaviors at timet.Based on the model setting, the values ofstandatare taken from corresponding state spaceSand behavior setAaccording to the above rules. The Q-learning algorithm generally uses ε greedy strategy to update strategy selection. In order to combine the Q-learning algorithm with the evolutionary game model, based on the consideration of the bounded rationality states of different individuals,the traditional ε greedy strategy is improved to obtain a bounded rationality evolutionary game strategy. The principle is shown in Fig. 8.

Fig. 8 Bounded rationality evolutionary game strategy

Under this strategy, all behaviors are selected with a non-zero probability ε. Due to the bounded rationality of an individual agent, it does not always learn from other individuals through the comparison ofQvalue in the selection of strategies. Therefore, in the evolutionary game strategy of bounded rationality, the individual randomly chooses the strategy in the next round of game with the probability of ε and compares theQvalue with that of other four random agents with the probability of 1-ε. Finally, the strategy with the largestQvalue is chosen as the strategy in the next round of game.

In the single-task and multi-decision evolutionary game model based on multi-agent and reinforcement learning,an individual agent selects the strategy used in each round of the game through the Q-learning algorithm. According to the evolutionary game rules, the state of the reinforcement learning algorithm is set as a three-person strategy combination. According to Table 1, the action set or strategy set isA={<Co Co Co>, <Co Co D>, <Co Co L>,<Co D D>, <Co L L>, <Co D L>, <Co Co Co>, <D D D>,<D L L>, <D D L>, <L L L>}, ten types in total. There is no difference of the order among strategy combinations.For example, the game rules used byCoCoD,CoDCo,andDCoCoare the same. The state space isS={0.96,0.98,1.1,1.02,1.04,1.06}, six types in total. The state space refers to the different values of β. The reward is determined by (14) and (15). The steps of the evolutionary game process are provided as follows:

Step 1Initialize theQvalue table.

Step 2Play a non-repeating three-person random game.

Step 3Calculate the rewards of all individual agents according to (14) and (15).

Step 4Select the strategy for the next round of games based on the bounded rationality evolutionary game strategy.

Step 5UpdateQtable according to (17).

Step 6Repeat Steps 2-5. Continue to play a new round of games until the specified number of rounds and then stop the game process.

The simulation is performed with the improved model under the reinforcement learning algorithm. Six values of β(0.96, 0.98, 1, 1.02, 1.04 and 1.06) are simulated respectively. The number of individual agents in the group is 500 and the reward-cost ratio is 3. The evolutionary game involves 1 000 rounds and the simulation is repeated 300 times.

The error ofQvalue under different β values can converge. In one simulation, the change ofQvalue error under the β value of 1 is shown in Fig. 9.

Fig. 9 Error of Q value

It can be seen from Fig. 9 that theQvalue error firstly fluctuates greatly, then gradually decreases and finally becomes stable with the increase in the rounds of the evolutionary game.

In a simulation of the evolutionary game under theβ value of 0.96, Fig. 10 shows the proportions of individual agents choosing different strategies in the group. It can be seen that in the evolutionary game process, the proportions of individual agents choosing each strategy fluctuate slightly. The proportion of cooperation strategies shows a clear upward trend, whereas the proportions of competition and inaction strategies show significant downward trends. The strategy with the highest proportion is the cooperation strategy and the strategy with the lowest proportion is the inaction strategy. The above results show that the introduction of reinforcement learning algorithms can effectively guide the group to evolve towards the direction of cooperation. In order to clearly show the proportion of individual agents choosing each strategy in the stable evolutionary game process under different values of negative feedback factor β, the average of the proportion of individual agents choosing each strategy in the last 100 rounds of the evolutionary game is used as the final evolutionary game stability result.

The proportions of individual agents choosing each strategy is shown in Fig. 11. It can be seen from Fig. 11(a)and Fig. 11(b) that under different β values, the proportion of individual agents choosing the cooperation strategy is significantly higher than those of individual agents choosing the competition and inaction strategies when the evolutionary game is stable. Moreover, the stable results of the proportions of individual agents choosing different strategies are not significantly different. The differences among the proportions of the same strategy under different β values are smaller after the reinforcement learning algorithm is introduced. The proportion of individual agents choosing the competition strategy tends to decrease as the value of β increases. When the evolutionary game is stable and β is 0.96, the proportion of individual agents choosing the competition strategy is the highest.When β is 0.96, the improved model allows the best guiding effect on the group evolutionary game. Whenβ is 1.06, the proportion of individual agents choosing the competition strategy is the lowest.

Fig. 10 Proportions of individual agents choosing different strategies in the evolutionary game

Fig. 11 Proportions of individual agents choosing each strategy under different β values in the stable evolutionary game process

When the value of β is 0.96, the evolutionary game effect of the model is the best. The evolutionary game results under the fixed selection value are compared with those based on the Nash-Qlearning algorithm, the Monte Carlo method and the genetic algorithm. All the algorithms have 1 000 rounds of evolutionary games and are repeated 300 times. Among them, the genetic algorithm takes every 100 rounds as a generation, and the game learning at the end of each generation adopts the classic uniform crossover operation of the genetic algorithm. When the evolutionary game of different methods is stable,the proportion of each strategy is obtained (see Table 2).

Table 2 Percentage of each strategy

It can be seen from Table 2 that the algorithm proposed in this paper has the best effect on the evolutionary game of the model since its cooperation strategy accounts for the highest proportion.

According to the previous analysis, the trend of the intelligence level of the group is related to the trend of the proportion of individual agents choosing the competition strategy. In the evolutionary game under the sameβ value, the proportion of individual agents choosing the competition strategy shows the increasing trend, so the group intelligence level also increases. In order to clearly show the changes in the group intelligence level under different β values, the group intelligence level in the initial stage of the evolution is compared with that in the final stable stage. The average value of the group intelligence level in the first 100 rounds is taken as the result of the starting stage and the average value of the last 100 rounds is taken as the final stable result. The group intelligence level is shown in Fig. 12.

Fig. 12 Group intelligence level in the initial stage and end of the evolutionary game under different values of β

It can be seen from Fig. 12 that under different β values, the group intelligence level in the stable stage of the evolutionary game is significantly higher than that in the initial stage of the evolutionary game. The group intelligence level is the highest under the β value of 0.96 and the lowest under the β value of 1.06.

The simulation results show that the single-task multidecision evolutionary game model based on multi-agent and reinforcement learning can effectively guide the group’s evolutionary game direction to evolve towards the ideal situation under the different effects of the negative feedback tax penalty mechanism. The improved model improves the group intelligence level and promotes the completion of the task. In reality, if the group performs an evolutionary game according to this model, the formulator of the income rules may not pay too much attention to relevant parameters in the negative feedback tax penalty mechanism, which can effectively guide the evolutionary game effect of the group. When the task completion requirements are high, relevant parameters in the negative feedback tax penalty mechanism should be considered in the formulation of reward rules.

The above simulation results are based on a group composed of a fixed number of people. In order to study the effect of the size of the group on the simulation results under different feedback factor values, 1 000 rounds of evolutionary game simulation experiments in the range of [100,500] are performed and repeated 500 times. Negative feedback factor β is respectively set to be 0.96, 0.98,1, 1.02, 1.04, and 1.06 and the reward-cost ratio is 3. The average of the last 100 evolutionary game results is used as the stable evolution result and 500 repeated experimental results are averaged. Fig. 13 shows the changes in the proportion of individual agents choosing each strategy as well as the average individual intelligence level in the stable evolution under different values of the feedback factor β. In order to more clearly display the changes under different reward-cost ratios, partial enlarged views are added in Fig. 13(b), Fig. 13(d), Fig. 13(f), and Fig. 13(h).

Fig. 13 Variations of the proportions of individual agents choosing each strategy and the average individual intelligence level with group size under different β values

Fig. 13(a)-Fig. 13(f) show that under different β values, the proportion of individual agents choosing the cooperation strategy in the stable evolution increases slightly as the group size increases. The proportions of individual agents choosing the competition strategy and the inaction strategy show downward trends. With the increase in the group size, the increasing or decreasing trend of the proportion of each strategy increases and the advantages of cooperation strategies become more significant. When the group size increases to 300 or more, the proportions of individual agents choosing different strategies remain stable. The further increase in the number of the individual agents will no longer affect the final proportion of each strategy. The dominant strategies under different β values are cooperation strategies and the strategies with the smallest proportion are inaction strategies. As the value of β increases, the proportion of cooperation strategies decreases, whereas the proportions of competition and inaction strategies increase. When the evolutionary game is stable and β is 0.96, individual agents choosing the cooperation strategy accounts for the largest proportion. When β is 0.96, the improved model shows the best guiding performance in the group evolution game. When β is 1.06, the proportion of individual agents choosing the cooperation strategy is the lowest. It can be seen from Fig. 13(g) and Fig. 13(h) that when the evolution is stable, with the increase in the group size, the average individual intelligence level increases slightly and remains stable. When the group size increases to above 300, the average individual intelligence level is no longer affected by the group size. When the evolutionary game is stable and the value of β is 0.96, the individual intelligence level is the highest and the rules of the game at this time have the best effect on the group’s cooperation and guidance.

Fig. 14 shows the variations of the penalty tax rate in the evolutionary game process with the group size in the stable evolution under different β values.

Fig. 14 Variation of tax rate levels with group size under different β values

As shown in Fig. 14, as the value of β increases, the penalty tax rate also increases and its value slightly decreases with the increase in the group size. If more people actually participate in the task, the value of the penalty tax rate decreases, but the effect of the decrease is relatively weak.

As shown in Fig. 13 and Fig. 14, the increase in the penalty tax rate increases the proportion of individual agents choosing the cooperation strategies and the average individual intelligence level, and the change in the group size has a limited impact on the level of the penalty tax rate and the proportion of individual agents choosing each strategy. With the increase of the group size, the evolutionary game direction of the group can be better guided towards the ideal situation, but the guiding effect reaches its limit when the group size increases to a certain degree.

6. Conclusions

In this paper, a multi-decision evolutionary game task is constructed to study the evolution rules of groups in the game process. After introducing the concept of multiagents into the evolutionary game process, a tax rate penalty mechanism with negative feedback characteristics is proposed to guide the selection of individual strategies in the group. In addition, a calculation method of the group intelligence level is defined to evaluate the result of the group evolution game. The simulation results show that the value of the negative feedback factor determines the strategy selection result of individual agents to a certain degree. When the value of the negative feedback factorβ is 1.06, it can effectively increase the proportion of individual agents choosing cooperation strategies as well as the group intelligence level and guide the group to evolve towards the ideal evolution direction. The tax rate punishment mechanism has limited guidance in the evolutionary game process. Most of the values of negative feedback factor β cannot effectively improve the evolutionary game results of the group. To solve this problem, this paper proposes a single-task multi-decision evolutionary game model based on multi-agent and reinforcement learning. The model combines multi-agents with Q-learning algorithms in the process of evolutionary games, improves the selection strategy of Q-learning, and proposes a bounded rationality evolutionary game strategy. This learning strategy not only reflects the rules of evolutionary games, but also takes into account the bounded rationality of individual agents. The simulation results of the model show that in the stable stage of the evolutionary game, different values of the negative feedback factor βcan guide the evolutionary game to evolve towards the cooperation direction and improve the group intelligence level. The simulation results also confirm the impact of the group size on the model. The increase in the group size can increase the proportion of individual agents choosing cooperation strategies as well as the group intelligence level to a certain degree. However, when it increases to a certain scale, it will no longer affect the results of evolutionary games.

In summary, the model proposed in this paper can effectively guide the evolution direction of the group in the single-task multi-decision game and explores the penalty tax rate and the group size in the evolutionary game.


登录APP查看全文