ラベル neural network の投稿を表示しています。 すべての投稿を表示
ラベル neural network の投稿を表示しています。 すべての投稿を表示

2021年9月25日土曜日

Running Q-Learning on your palm (Part3)

Japanese summary 本シリーズのPart1とPart2では、スマホ向けのQ-Learningアプリを開発し、それを簡単な例(直線の廊下でロボットが宝石を得る)に適用しました。今回は、このアプリを改訂し、2次元グリッドでロボットが行動できるようにしました。そして、グリッドサイズが大きくなるにしたがい、これまでのQ-tableを保持してQ値を更新する方法は、メモリ量と処理量の急増により破綻することを確認しました。それに変わる有望な方法として、Neural Networkの利用を検討し、それを(スマホではなく)PC上のPythonで実現した結果を示します。

Abstract
In Part 1 and Part 2 of this series, I developed a Q-Learning app for smartphones and applied it to a simple example (a robot gets a gem in a straight corridor). This time, I revised this app so that the robot can act on the 2D grid. However, as the grid size increases, the traditional method of holding the Q-table and updating the Q value becomes difficult due to the increase in memory capacity and processing volume. As a promising alternative, I considered using a neural network. And here's the result of doing that with Python on a PC (not a smartphone).

● Revised version of the Q-Learning app
In the revised version of the smartphone app, as shown in Fig. 1(a), the robot is trained to reach the gem while avoiding barriers on the 4x4 grid. The Q-Learning algorithm is basically the same as last time. There is one gem and one barrier, and their positions change randomly for each play (every episode). The robot also starts from a random position. After sufficient training, the robot can always reach the gem in the shortest route. On the other hand, if not well trained, the robot often gets lost and hits a wall, as shown in Fig. 1 (b).


● Memory capacity required for Q-Learning
The size of the Q-table required for this learning can be calculated according to the grid size and the number of gems and barriers. See Fig.2. The number of Q-table entries (i.e., the number of keys) is the total number of possible states, which in Case 1 (4x4) is 3,360. At this level, it can be held sufficiently even on a smartphone, and the amount of calculation is within an acceptable range. However, in Case2 (5x5), the total number of states increases sharply to over 6,000,000, even though only one gem and one barrier have been added. In this situation, regardless of whether it is a smartphone or a PC, processing is almost impossible due to both the amount of memory and the amount of calculation.

● Calculate Q-values with neural network (without holding Q-table)
For cases like Case2 above, you can think of a way to calculate the required Q-value with a neural network without holding the Q-table. To do this, transform the Q-value update formula for the neural network, as shown in Fig.3. This makes it possible to compare output and target  ([1]). It can be used to solve this problem with common supervised machine learning. This machine learning iteration allows the output to be closer to the target and, as a result, the Q-value to be closer to the exact value.

Fig.4 clearly shows how to use this neural network in the case of Case1. Note that in this example, the action "W (west)" is taken in the current state S. In this way, one learning is done only for one action in one state. This learning should be repeated for as many actions as possible, in as many states as possible.

● Calculation example of Q-value by neural network
I implemented a learning method using such a neural network in Python and executed it on a PC. This program is based on the Python program (using Tensorflow / Keras) published by Dr. Makoto Ito in reference [1]. Fig.5 shows the learning process for Case1 (4x4). It shows the situation where 10000 episodes were randomly trained. In the upper graph, the average sum of rewards per episode has reached about 0.8. On the other hand, when the neural network is not used, as shown in the figure on the right of Fig. 1(a) (although the characters are small and difficult to see), it is 0.8305, so both results are almost the same. The lower graph shows that the average number of steps a robot takes to reach the gem is about 2.9. This value is also valid considering the situation in Fig.1.


I have omitted the details, but in the case of Case2 (5x5), I was able to train well with this neural network as well. It took only about 3 minutes to run on a general PC, so I was able to confirm the usefulness of this method. This time I've only used the most basic neural networks, but for more complex problems (for example, if you need to remember the location of an object), you may need other neural networks such as LSTMs.

Acknowledgments
I was able to create a Q-value calculation program using a neural network by referring to the Python program published in the reference [1]. I would like to thank Dr. Makoto Ito, the author of this article.

References
[1] Makoto Ito's Blog Article: M-note Program and Electronic Work (in Japanese)
     http://itoshi.main.jp/tech/

2019年3月19日火曜日

Neural Networkでエージェントの頭脳を作る(agent-based modeling)

本稿は、エージェント指向とニューラルネットワークの連携に関する記事です。
(図は細かいので、クリック拡大してご覧下さい。)

 Agent-Basedモデリングの有名な例のひとつにWolf-Sheep Predatorがあります。ここではNetLogoによるモデリングを扱います。これによって、下記のような動作をする3者の個体数がどのように推移するかを観測できます。
  1. 地面(茶色)には、所々にgrass(草、緑色)が生えている。
  2. grassは、sheep(羊、白色)に食べられるが、一定時間後には再生する。
  3. sheepはgrassを食べてエネルギーを得る。
  4. wolf(狼、黒色)はsheepを食べてエネルギーを得る。
  5. sheepとwolfは、動くたびに一定のエネルギーを消費する。
  6. sheepとwolfは、エネルギーがゼロになると死滅する。
  7. sheepとwolfは、一定の確率で子を産む。
 従来のモデルでは、多数のwolfとsheepがそれぞれランダムに(緩やかにある方向へ)移動していました。しかし、下記の論文[1]には、wolfとsheepに頭脳(Neural Network)を持たせて、周りの状況に応じた適切な方向へ移動させる改良がなされています。Neural Network自体は、Wolf-Sheep Predatorとは独立した別モデルとして作られています。両者を連携させるために、NetLogo6以降で導入されたLevelSpaceという拡張機能[2]を使っています。

 各個体のNeural Networkは、Fig.1に示すように、9ノードからなる入力層、9ノードからなる中間層、3ノードからなる出力層で構成されます。この論文と共に、NetLogoのソースプログラムも公開されているのです。しかし、コメントがなく、ラムダ式をふんだんに使っており、ちょっと難解でした。本稿では、当方でそれを解読した結果に基づいて述べます。


 もう少し詳しくみてみましょう。Fig.2はあるsheep(sheep #3)のある時点での頭脳を示しています。入力層で"1"となってノードは(横に赤字項目がある)、周囲からそこへ刺激があったことを意味します。すなわち、以下の3項目です。
  • 左方向(一定範囲のcone vision内)にgrassを見た。
  • 右方向(一定範囲のcone vision内)にgrassを見た。
  • 正面方向(一定範囲のcone vision内)にwolfを見た。
 最終段の出力層は、softmaxによる出力です。「左へ向かう」、「直進する」、「右へ向かう」のうち、最も強い「右へ向かう go right」という判断がなされました。このsheepにとっては、左右方向に草があり、正面方向にwolfを見たのですから、これは確かに、妥当な判断と言えます。


 ここまでは、少ない個体数で想定しましたが、次に、sheepが100匹、wolfが50匹という少し大きな世界をモデリングします。この場合は、150個の独立したニューラルネットワークが同時に動きます。(途中で生死により増減します。)それらすべてを表示するのは現実的はありませんので、Neural Networkの表示を消してシミュレーションを実行しました。その様子をFig.3に示します。図の中のグラフから分かるように、この世界では、周期的に変化はしますが、wolf、sheep、grassの個体数のバランスが長期的に保たれているように見えます。


(注1)基本的には、Fig.2の例のように、出力層のsoftmax値の最大値が移動方向を決めるが、ある確率でそれを揺らしている。すなわち、いつもそのとおりの方向へ移動するとは限らない。(例えば、あるwolfが、右方向にsheepを見てそちらへ向かっても、次の時刻には、sheepはwolfの移動方向と離れてしまうかも知れないですから。)

(注2)ニューラルネットワークの辺の重みは、別途、誤差逆伝播によって学習されるはずです。このモデリングでは、学習済みになった重みを使います。また、wolfとsheepが産んだ子供には、親のニューラルネットワークが引き継がれますが、その後に、僅かな変異(辺の重みの修正)を起こせるように設計されています。

参考資料
[1] Bryan Head, Arthur Hjorth, Corey Brady, and Uri Wilensky, "EVOLVING AGENT COGNITION WITH NETLOGO LEVELSPACE", Proceedings of the 2015 Winter Simulation Conference, pp.3122 - 3123.
[2] Extensions LevelSpace, https://ccl.northwestern.edu/netlogo/docs/

2018年6月14日木曜日

BackPropagation in Neural Network with an Example (XOR) Using Multiple Micro:bits

In this article I would like to attempt Backpropagation in a neural network using micro:bit.

Back propagation in a neural network made with micro:bits

In the previous article, I used micro:bit to construct a neural network and took up the recognition of XOR using it.

http://sparse-dense.blogspot.com/2018/06/microbittwo-layer-perceptronxor.html

At that time, since we used the neural network weights completely learned about z = XOR (x, y), the value of z for any input pair (x, y) could be calculated with only one forward propagation.

As before, I use micro:bit neural network again. But at this time, the Backpropagation which is the foundation of Deep Learning is applied. In general, it takes a long time to learn in a neural network, since it requires repetition of forward propagation and backward propagation many times. So, here we use some good (but not complete) weights that we got before from another projects, as initial weight values. Starting from that, we should complete weights update within a small number of iterations, even on the micro:bit neural network. Please check the actual operation with the following video:



This video is also published on Youtube:
https://youtu.be/tsYr01lQ_HY

Weight updates to recognize XOR are required for the four cases of input (1, 1), (1, 0), (0, 1), (0, 0). Successively for all these cases, updates should be repeated until the error (difference between the answer and the evaluated value z) becomes smaller than the tolerance (here, 0.1). However, for the sake of simplicity, we only show weight updates for the first case x = 1, y = 1 in this demonstration. It converged after 5 iterations.

Please note that the neurons (micro:bit) of the Hidden layer and Output layer display eastward arrows or westward arrows depending on whether they are in Forward or in Bacward. Also, after the Backward is over, button A on micro:bit is pressed before Forward starts. This is to ensure the proper timing between receive and send on micro:bit radio.

2018年6月4日月曜日

Consturuct a neural network (multilayer perceptrons) using micro:bit

notice:
This article treats only forward propagation, whereas, backward propagation is also demonstrated in another article. -> Please see here.

Let's learn the basics of neural networks using micro:bit. Understanding deepens by actually programming. For neural network learning to solve various tasks, back propagation is generally required. Learning with this back propagation requires considerably long time calculations. It is not realistic to do this with such a tiny microbit with low computing power, even though it is not impossible. Therefore, here we will try forward calculation only, using the edge weights of the already learned neural network. Even so, you can experience some of the important points of the neural network.

Neural network with micro:bit

 this illustrates evaluation of z = XOR(1,1) = 0

One thing to notice here is that one microbit plays a role of one neuron. By doing so, you can have an image that is close to real neurons. Here, XOR (Exclusive OR) is the problem. z = XOR (x, y) shows 1 only if either x or y is exclusively 1. Otherwise the value of z is 0. In the microbit, JavaScript and MicroPython can be used, but here we use MicroPython. The reason is that microbit JavaScript currently does not support calculation of float type data. Float type operation is mandatory for calculation of Neural network. Also, neurons (microbit) exchange signals with each other. For that we use send / receive on radio. This is not Bluetooth. Additionally, as the activation function, we use the sigmoid function.

The weights of the edge used above are the results of being learned by the following NetLogo program that is based on the following one:
The red line denotes a negative value, the blue line denotes a positive value, and the thickness shows the magnitude of the absolute value.

Weights used in the above example were obtained by this NetLogo  simulation

Please see the video for an actual operation example.
https://youtu.be/PPUcsXgCnZ4

In the first half, z = XOR (x, y) is calculated where x = 1, y = 1, and then the value 0.0178... was obtained, so that "0" was finally displayed. On the other hand, the second half is the case of x = 1 and y = 0. Now that 0.9852... has been obtained, "1" was displayed as the final result of z.