Introductory Math4AI [English Lecture] K-MOOC for Public!!


    <Introductory Mathematics for Artificial Intelligence>


    일반인을 위한 K-MOOC


     <인공지능을 위한 기초수학 입문>


    그림입니다.
원본 그림의 이름: mem000015b4a196.jpg
원본 그림의 크기: 가로 3019pixel, 세로 321pixel
사진 찍은 날짜: 2019년 12월 12일 오후 6:43
프로그램 이름 : Microsoft Windows Photo Viewer 6.3.9600.17415


    http://matrix.skku.ac.kr/math4ai-intro/

     

    인공지능을 위한 기초수학 입문 [실습실]


    01주차 http://matrix.skku.ac.kr/math4ai-intro/W1/

    02주차 http://matrix.skku.ac.kr/math4ai-intro/W2/

    03주차 http://matrix.skku.ac.kr/math4ai-intro/W3/

    04주차 http://matrix.skku.ac.kr/math4ai-intro/W4/

    05주차 http://matrix.skku.ac.kr/math4ai-intro/W5/

    06주차 http://matrix.skku.ac.kr/math4ai-intro/W6/

    07주차 http://matrix.skku.ac.kr/math4ai-intro/W7/

    08주차 http://matrix.skku.ac.kr/math4ai-intro/W8/

    09주차 http://matrix.skku.ac.kr/math4ai-intro/W9/

    10주차 http://matrix.skku.ac.kr/math4ai-intro/W10/

    11주차 http://matrix.skku.ac.kr/math4ai-intro/W11/

    12주차 http://matrix.skku.ac.kr/math4ai-intro/W12/

    13주차 http://matrix.skku.ac.kr/math4ai-intro/W13/

    14주차 http://matrix.skku.ac.kr/math4ai-intro/W14/


     

    ‥  목차   Contents

     


    서 문                                                                       8


    I.  인공지능에 필요한 기초수학 (1주차)     14

      1. 함수 그래프와 방정식의 해  14

      

    II.  인공지능과 행렬 (2-6주차)        25

      2.  데이터와 행렬 25

      3.  데이터의 분류 38

      4.  선형연립방정식의 해집합 48

      5. 정사영과 최소제곱문제 58

      6.  행렬분해 (특잇값 분해) 65

     - 과제 -  74


    III.  인공지능과 최적해 (미분) (7-9주차)     75

      7. 극한과 도함수 75

      8. 극대, 극소, 최대, 최소  88

      9. 경사하강법, 최소제곱문제의 해 95

      - 과제 -  107


    IV.  인공지능과 통계 (10-11주차)     108

      10. 순열, 조합, 확률, 확률변수, 확률분포, 베이지안(Bayesian)  108

      11. 기댓값, 분산, 공분산, 상관계수, 공분산 행렬 122

      - 과제 -  133


    V.  주성분 분석과 인공신경망 (12-14주차)   134

      12. 주성분 분석(Principal Component Analysis) 134

      13. 인공신경망 (Artificial Neural Network) 145

      14. MNIST 데이터 숫자인식 실습 157

      - 과제 -  166


    VI. [읽을거리, 참고문헌, 기타] 수학과 코딩                           168


    참고 자료 [Math4AI]  http://matrix.skku.ac.kr/math4ai/

    교재 다운(Down), Math &Coding [Down] 


    Part Ⅰ. 행렬과 데이터분석            http://matrix.skku.ac.kr/math4ai/part1/ 

    Part Ⅱ. 다변수 미적분학과 최적화     http://matrix.skku.ac.kr/math4ai/part2/ 

    Part Ⅲ. 확률통계와 빅데이터          http://matrix.skku.ac.kr/math4ai/part3/ 

    Part Ⅳ. 빅데이터와 인공지능          http://matrix.skku.ac.kr/math4ai/part4/

                     http://matrix.skku.ac.kr/2020-math4AI-final-pbl/


     

    1. 홍보 (Welcome to Math4AI)


    Nice to meet you, everyone.


    I am Sang-Gu LEE, a professor at Sungkyunkwan University, who will teach K-MOOC <Introductory Math for Artificial Intelligence>.


    Artificial intelligence came very close to our life already.


    For example, Netflix, Amazon's product recommendation system, AI voice assistant 'Siri', AlphaGo Master who won all the best Go players, and NEON and Bally which was first introduced at CES 2020.


    In addition, voice recognition, face recognition, self-driving car, understanding and generation of natural language are all related to AI.


    As emphasized in the <Mathematics Changes the World> series, which won the Grand Prize in the Korean Journalist Award at Citi-Group, mathematical theory was a key factor in the development of AI.


    In the near future, artificial intelligence robots will be working for humans in factories, roads, and homes. As a result, humans may not have to go to work and lose our jobs.


    But we must not be afraid of AI.


    Instead, we need to understand the basic principles of what artificial intelligence is and how it works.

     

    In this lecture, we will explain all the basic mathematical knowledge that is necessary to understand AI.


    To help you understand, I offer free web contents, free cyber lab using Python-based free cloud computing tools.


    And we will use electronic textbooks and lecture note to help you to understand a basic math for AI.


    Welcome and let's be together for the next 14 weeks.



    2. AI 란 무엇인가? (What is AI?)

    0-2.


    How are you?  Welcome to K-MOOC <Introductory Mathematics for Artificial Intelligence> course.


    When I was asked to teach Math for AI, I wanted to know more about AI. And the book <Practical AI for Dummies> by Christian Hammond was recommended to me. Since it was very easy to read, I translated it. Now I am going to introduce the main contents of this book.


    Artificial intelligence come to us very closely already. AI now cooks and wash dishes for us, gives answers to questions and schedules for us.


    We now need to understand AI and how it works. So we set the goal of this course is to learn and practice the contents of linear algebra, statistics (especially PCA, Artificial neural network, etc.) that is needed to understand artificial intelligence.


    <AI for Dummies> covers:

    Definition of Artificial Intelligence; It is the process of recognizing, deducing, and operating a situation, extracting information from data, Deep Learning, and Prediction Analysis, especially natural language processing etc,


    If we want to know the details, I may refer a video which I prepared.


    It seems that the age of robots is already nearby us.


    In the past, we saw robots in movies like Terminator, Matrix, and iRobot.


    Now robots drive, and work for humans at home.


    Therefore, it is important for us to understand how AI works. This book <Practical AI for Dummies> introduces that.


    Already, robots are going beyond human abilities in quiz games, Go- games, and Internet games. In the last chapter of the book, 10 tips that we must know for the future in the age of robots was introduced.


    The outline of this book is as follows; the definition of AI, the story of the 1956 Dartmouth Conference, where artificial intelligence was first introduced, five reasons that AI is reemerging recently, recommendation process in AI, examples of AI, reasoning, and behavioral steps are not magic but they are functional processes by mathematical algorithms.


      Let's take an example. When you say "I want Pizza" to Siri, AI confirms that there is no such term like "cooking," and judges that the speaker is looking for a pizza place to eat. And AI uses GPS information to find nearby pizza restaurants, ranking them according to proximity, reputation, price, etc. Also AI refers to your records of restaurants visit, then recommends a high-scored restaurant recommended by other people. As a result, AI generates a sentence “There is a pizza restaurant called Gino's Pizza about three blocks from here.” and reads it in a voice.


    The book introduces 'The Handwritten Number Recognition System Using MNIST Data Set'. It explains "recognizing that a given letter from the data is a number close to 7" and "the code that implements the artificial neural network to recognize which number this handwritten letter is close to." And using these codes, we can do the process of increasing the accuracy of the neural network by changing the number of data, layers, and nodes.


    The book introduces the next sentence by Red Headed League author about 'how to choose the necessary information from the data when drawing conclusions about the process of extracting information using a big data'.


    "At first I thought you were a genius, but now I know that anyone with observation and analysis could do the same thing.”


    Here is another example, it explains the operational process of the purchasing system.


    If there is a record of the books you've purchased so far, artificial intelligence selects people with a similar purchasing history to you, checks out the books that they bought recently and books that you haven't bought yet, then recommends a book by email or SNS. It also introduces the operating principles of marriage brokerage firms.


    It also introduces deep learning and smart devices, explains neural networks, and introduces important mathematical concepts.

     

    You will see the details in https://youtu.be/F1HNFGAMhro. You will see 'what artificial intelligence is' by watching the video and will also see what math contents are essential to understand AI.


    Then we will be ready for this class of Introductory Math for AI.


    If there is anything you want to know more, feel free to ask in QnA board.


    With the above background, we will be ready to start this K-MOOC <Introductory Math for Artificial Intelligence>. Let’s start.


    3. AI & Mathematics!! 

    AI 0-1.


    Hi, everyone! Today we will have an introduction session for <Introductory Mathematics for Artificial Intelligence>.


    You may think artificial intelligence has nothing to do with you. However, AI is already a key driver of the Fourth Industrial Revolution and brings many changes in industry and social structure through technological innovation.


    We are already living with artificial intelligence, for example, FinTech. Healthcare, Drones, Self-driving cars, Smart homes, AR, VR, etc.


    Artificial intelligence with Data and Networks, is being developed into various types and is driving innovations in all industries.


    Artificial intelligence has been used for a long time in Korea. In 1991, Lucky-Goldstar (now LG) promoted and sold <Artificial Intelligence Chaos Washing Machine>. Recently, a robot vacuum cleaner Code-zero R9 is cleaning the house on behalf of humans. At CES 2020, I saw an AI robot chef that cooks and wash dishes for us.


    In these days, AI environment has enormous potential. Killer robots can threaten the world, Supercom <WATSON> competes with knowledge at the television quiz show <Jeopardy!> and AI secretary Siri tells us where to eat.


    When the age of robots come, we thought that the emotionless robots (which look like Arnold Schwarzenegger) might attack and kill humans. And we think that superintelligent robots will dominate the Earth, or we were afraid of the situation as we saw in the movie “Matrix”.


    We now see that many of us fear AI.


    In the movie <i-robot> one said that 'AI robots recognize me and lie to me and wink at me. I'm afraid I'll be living in a daze without thinking until then.' In other words, what we really need to fear is to live without knowing ‘What Artificial Intelligence really is’ for the rest of our lives.



    The concept of artificial intelligence was first introduced at the Dartmouth Conference in 1956, and has developed steadily so far.


    In 2011, the WATSON won the Jeopardy Show, and in March 2016, AlphaGo competed against legendary Go player Mr. Lee Sedol. Since then, AI began to get closer to us. Currently, there are AlphaGo Master and AlphaGo Zero.


    The development of AI has two dark periods. However, in early 2000, deep learning algorithms have been developed, massive data is being accumulated and computing power is evolving. Accordingly, the AI era is growing in every way.


    Artificial intelligence is everywhere in our life. AI talks with humans, recommends items to buy, advises finance, solves quizzes, quickly and accurately identifies the patients' disease, finds court precedents, and wons Go. In particular, AlphaGo won Lee Se-dol, AlphaGo Master won Ke Jie, Alpha Star competes with Starcraft 2 and  Alpha Zero learns the rules of the game by himself and started to win many games.


    Google and Facebook also introduced voice recognition, face recognition.  Self-driving cars depends on deep learning and the AI system started to understand and create Natural Languages.


    An emotionless robot works, selects what we want to buy, and tells us where we like to go. As we saw that AI is so good to do these things, we gradually became afraid of AI.


    Robots work for humans in factories, roads, and homes.


    Sooner or later, we will take self-driving Google cars to work, read a newspaper without any traffic jams. In some day, we even may not need to go to work anymore.


    Recently, I was introduced to a book that helps to reduce this fear named <Practical AI for Dummies by Kristian Hammond>.

    I translated this book, explained it to my students, and recorded it. You may like to see it in https://youtube/F1HNFGAMhro.


    Why is AI emphasized in these days?


    Artificial intelligence is a concept that computers behave like a human brain, recognize images, and perform tasks such as self-driving.

     

    Therefore, machine learning is a subset of artificial intelligence.


    The key we need to know is that "AI is not a Magic."


    AI is an application of algorithms, or collections of mathematical algorithms that are operated based on data with a processing power.


    When Amazon recommends you a book, it is important for you to understand the AI system behind this process.


    Amazon collects informations on your buying patterns, checks who you are, finds people who are similar to you, and recommends new products based on their shopping histories.


    There are five key reasons why AI is emphasized now.


    First, the computer's performance has been improved.


    Second, the big data is exploding.


    Third, AI handles well a certain problem. For example, systems such as Siri and Cortana, have been developed to choose very specific words from what people say.


    Fourth, the computer system has changed into a self-learning environment.


    Finally, we now realized that robots don't need to have reasons unlike humans and hence started to develop alternative reasoning models. Therefore, we have admitted that machines think like machines and began to study and apply artificial intelligence in a way that maximizes and optimizes the way of Robot thinking. As a result, AI started to change our lives drastically.


    Therefore, we must understand what the basic theory of reasoning in AI systems, and how AI systems can find and provide answers that humans need. And we should know mathematical knowledges are needed to understand this AI system.


    For example, consider a self-driving car. These are photos of self-driving cars taken at CES 2020. If we tell our destination to the cars, then the car find its own direction and starts driving. The drone identifies a dangerous object for the car.


    While AI is driving, we can talk comfortably in the car, prepare for a meeting, and arrive safely at our destination. 


    As such, artificial intelligence is affecting almost every part of our lives.


    The important thing is to understand that this artificial intelligence is a <Mathematics>, not just a <Magic box>. In general, a people say, 'Artificial intelligence is complicated!'. But AI is actually a Math, so mathematicians can explain it very easily.


    For example, artificial neural networks are functions and matrices. Now it plays a major player in artificial intelligence.


    Large hospitals have medical records of many patients, which show persons’ height, sleeping hours, exercise hours, calorie intake, weight, and blood pressure, etc.


    If we give the inputs such as sleeping hours, exercise hours, calorie intake, and height into AI, then AI gives us an output such as a normal weight or a blood pressure. This is a process of the function . 


    If both Input and Output are vectors, then the matrix acts as a function. Therefore, if we are able to find this matrix, it is possible to predict that 'people with certain inputs will have certain outputs'.


    This is just a linear algebra.


    In the past, we try to build up a Mathematical model to be used for our data. But now when we have enough input data and answers, a Machine finds rules, functions, or matrices from those data on its own.


    In particular, artificial neural networks are modeled on neurons, the basic unit of a nervous system. The human brain has about 100 billion neurons and performs well tasks like image recognition. The .\PICture below is a real neuron and an artificial neuron.


    When the nerve is stimulated, it does not respond to small stimuli, but when the stimuli go beyond threshold value, it reacts with "Ahhhhh!"


    Sigmoid function is used for these neural responses. Artificial neural network is a model of this process of outputting results through the layers when a data is entered.


    Now let's explain this process mathematically.


    We can think of a function. When we give (input) x, (output) y=ax is given.


    Here, we use the symbol w  instead of a.


    If we give (input) x, then (output) y=wx is given, here x and 1 are input, and y=wx+b is an output, when x_1 and x_2 are input, y=(x_1w_1+x_2w_2)+b will be an output, and when x_1, x_2, x_3 are input,  y=(x_1w_1+x_2w_2+x_3w_3)+b is an output.


    As in the .\PICture, if there are several Inputs and Outputs, then it can be seen as the product of a matrix and a vector (from Part 1) to have the output.


    At this time, the bias is assumed to be zero for simplicity, and this w_ij is called the weight from the i-th node of the input layer to the j-th node of the output layer.


    It 4×1 vector is a input, then it is multiplied by a 2×4 matrix A, so a 2×1 vector is an output. In other words, computations in artificial neural network are just like a matrix operation.


    In mathematics, this is y=AX, but the notation of x^T A^T = y^T are more popular in artificial intelligence or statistics.


    That is, when a 1×4 vector [x_1 x_2 x_3 x_4] is a input, then a 1×2 vector [y_1, y_2] is an output which is obtained by multiplying the 4×2 matrix W = [w_ij].


    The x^T A^T = y^T is a transpose of y=AX.


    Therefore, it is important to find a matrix A.


    It can be done when we have enough data to find a matrix A. Using sufficient data, we gradually refine the function (Matrix) to obtain a matrix A.


    After that whatever Input is given, we can just multiply the matrix A to predict the Output.


    We now see this process is not a magic.


    This process uses mathematical and statistical techniques to improve the function(Matrix).


    The above process is called a Backpropagation algorithm.


    In other words, a back propagation is a process of gradually increasing the precision of the desired matrix by correcting errors while going back and forth between the hidden layers. 


    A Backpropagation is a process of minimizing the error function to improve the neural network. [Step 1] In finding a Least Squares solution, the Singular Value Decomposition is used. [Step 2] And the Gradient descent method is used as an algorithm to minimize error in this error function.


    We will learn mathematics that are needed to obtain this matrix A and complete our mathematical modeling.


    Let's take an example. When you say "I want Pizza" to Siri, AI confirms that there is no such term like "cooking," and judges that the speaker is looking for a pizza place to eat. And AI uses GPS information to find nearby pizza restaurants, ranking them according to proximity, reputation, price, etc. Also AI refers to your records of restaurants visit, then recommends a high-scored restaurant recommended by other people. As a result, AI generates a sentence “There is a pizza restaurant called Gino's Pizza about three blocks from here.” and reads it in a voice.


    Take another example. Let's explain how retailers recommend products for their consumers. First, AI identifies people who have similar purchasing propensity to you and checks what these people bought and shopped, then deduces these people's behavioral records in the following ways: You and Bob both like snowboard, kickboxing, and skydiving. Bob likes water skiing. Then you will like water skiing, too.


    If there are 270 million Bob-like people in the world, companies can predict very accurately what you like and what you may buy next in this way.


    Bob bought books A, B, C, D, and you bought them too. Recently Bob bought a book called <Data Science for Beginners>. Then, Amazon works and recommend you that "Don't you interested to buy a new book <Data Science for Beginners>?".


    In other words, AI analyzes the data, checks what people need, deduces what they will like, and constantly sends recommendation messages. You keep getting these messages, too.


    Like this, AI works according to a process that we can fully understand. AI is not a smart box that we don't know. In other words, AI is not a magic. AI is based on algorithms that operate using data and processing power.


    Finally, let's explain the English-French translation process.


    Computers will have billions of English-French sentences that express the same idea and learns the way of translating the two languages into each other by itself.


    When we enter a English sentence, computer finds a similar French sentence using the knowledge learned from those big data.


    For IBM Watson, 90 machines run in parallel processing. Therefore, they perform a lot of searches in a very short time and inform you the results in a very short period of time.


    Artificial intelligence is a very broad concept, machine learning is a subset of artificial intelligence, and the deep learning is a subset of machine learning. And recently a deep learning is a hot issue.


    As artificial intelligence affects life, many companies are choosing artificial intelligence.


    We can control artificial intelligence rather than artificial intelligence does it to us.


    To do so, we must help students understand math, statistics and coding step by step.


    Students should learn eigenvalue, eigenvector, diagonalization of matrices, SVD, PCA, local extremum points, and discrete mathematics in Mathematics subject, probability distribution, covariance matrix in Statistics subject, and some Coding experience using Python language and R language.


    If we do well, they will have the ability to understand and utilize artificial intelligence, and will be able to use artificial intelligence without fear and prepare for the future job.


    Artificial neural network is a key player in artificial intelligence.

    Artificial neural network produces a function that predicts output values from input values. At this time, if both input and output are vectors, the function is going to be a matrix.


    If we can find that matrix (or understand the process of finding the matrix), we can always estimate the proper output from the given data.


    If we understand that the matrix itself is an artificial neural network, it will be easy to understand <Introductory Mathematics for Artificial Intelligence> without fear.

     

    Therefore, we select and explain the knowledge required for artificial intelligence such as linear algebra, calculus and statistics.


    And we will use 'text Code' to do computations and simulations.


    In particular, we have prepared a textbook with code and hence we can practice problems directly in Python language and R language.


    As a result, we can handle functions with Code, so we can easily analyze the given data.

     

    We are going to explain how to connect 'Math and Coding' to maximize your ability to handle AI.


    I look forward to your pleasant learning of  <Introductory Mathematics for Artificial Intelligence>.  Thank you.

     


    4. Syllabus (강의계획서)


    안녕하십니까.

    K-MOOK 인공지능을 위한 기초수학 입문을 담당하게 된 성균관대학교 자연과학대 수학과 이상구교수입니다. 반갑습니다.

    이번 고교생과 일반인을 위한 K-MOOK introductory mathematics for artificial intelligence 즉, 인공지능을 위한 기초수학 입문.


    Welcome! May I introduce myself? 


    I am Prof. Sang-Gu LEE, at Sungkyunkwan University, who is in charge of the K-MOOC <Introductory Math for Artificial Intelligence(AI)> class for everyone.



    2020 경문사에서 발간된 인공지능을 위한 기초수학 입문 introductory mathematics for AI 교재를 활용할 것이며, 그 내용은 math4ai-intro 웹 주소에서 확인할 수 있도록 제공하였고,


    This class will use the textbook <Introductory Math for Artificial Intelligence, Kyungmun Books, 2020>.

    And my lecture note can be found in the following web address:

                    http://matrix.skku.ac.kr/math4ai-intro/  .


    목차는 Part 1. 인공지능에 필요한 기초수학, Part 2. 인공지능과 행렬, Part 3 인공지능과 최적해, Part 4, 인공지능과 통계, Part5. 주성분 분석과 인공신경망에 대해서 다루고 이 전체 내용을 14주에 걸쳐서 공부하겠습니다.


    The contents of the class are given as follows:

     

    Part 1. Fundamental Mathematics for AI

    Part 2.  AI and Matrix

    Part 3  AI and Optimization.

    Part 4,  AI and Statistics

    Part 5. Principal Component Analysis(PCA) and Artificial Neural Network(ANN)


    We will study these 5 parts in 14 weeks.



    평가방법은 출석, 참여, 10회에 걸친 퀴즈에서 40%, 중간고사 30% 그리고 보고서와 기말시험 30%, 총 100% 비율로 평가할 것입니다.


    The evaluation will be divided into 40% for attendance/participation/quizzes (10 times), 30% for Midterms, 30% for Reports and Final exam, for a total of 100%.


    인공지능을 위한 기초수학 입문 내용을 좀 더 자세히 설명 드리면, 첫주차에는 인공지능에 필요한 기초수학으로 다양한 함수의 그래프를 그리고 그것을 이용하여 방정식의 해를 구하는 방법을 복습합니다. 그 다음 인공지능과 행렬에서는, 5주에 걸쳐서 선형대수학 지식을 정리한 데이터와 행렬, 데이터의 분류, 선형연립방정식의 해집합, 정사영과 최소제곱문제, 행렬분해 특히 특잇값 분해를 다룹니다. 인공지능과 최적해에서는 미분, 도함수, 편도함수 내용을 3주에 걸쳐 다루면서 극대와 극소, 경사하강법, Gradient Descent method를 이용하여 최소제곱문제의 해를 구하는 방법을 이론 및 실제 실습을 통하여 학습할 것입니다. 그리고 2주에 걸쳐서 통계의 기초 지식인 counting method 순열, 조합, 확률, 확률변수, 확률분포 그리고 기댓값, 분산, 공분산, 공분산 행렬에 대해서 학습할 것입니다. 마지막으로 주성분 분석 PCA와 인공신경망에서는 Principla Component Analysis, Artificial Neural Network 그리고 MNIST 데이터 set을 이용한 숫자인식 이론과 실습을 하고 본 강좌를 마칠 계획입니다.


    Let's talk more about  <Introductory Math for Artificial Intelligence(AI)>.


    In the first week, we will review some mathematical fundamentals that are needed in the study of AI; For example, Plotting graphs of functions and Finding the solutions of equations.


    For the next five weeks in Part 2 <AI and Matrix>, we will study basic linear algebra; Data and matrices, Classification of data, System of linear equations, Projections, the Least squares problem, Matrix decompositions including the Singular Value Decomposition.


    For the next three weeks in Part 3 <AI and Optimization>, we will study basic Calculus; Derivative, Partial Derivative, Gradient, Local Extremum, Absolute minimum, Gradient Descent method. We will use the Gradient Descent method to find the solution of a least square problem.


    For the two weeks in Part 4 <AI and Statistics>, we will study basic statistics: permutation, combination, probability, probability variable, probability distribution, and expected values, variance, covariance, covariance matrix.


    Finally, we will use two weeks for Part 5 <PCA and ANN>, you will learn Principle Component Analysis, Artificial Neural Network and an ANN example using MNIST Data set.


    본 강좌를 효과적으로 운영하기 위하여 웹 컨텐츠를 제공하면서 동시에 사이버랩, matrix.skku.ac.kr/KOFAC 한국과학창의재단과 협조하여 만든 사이버랩을 이용하여 파이센 기반의 코드를 사용하면서 실제 계산과정을 획기적으로 줄여서 데이터들을 다루고 시뮬레이션 할 수 있도록 준비하였습니다.


    We will provide Web content and CyberLab for you to use to learn the content more effectively. http://matrix.skku.ac.kr/math4ai-intro/ and

    http://matrix.skku.ac.kr/KOFAC/ 


    You will be more confident for computing and simulating your data with the provided Python based Sage codes and R codes.


    그리고 본 강좌에서는 10회에 걸쳐 간단한 온라인 퀴즈를 제공할 것이고, 질의 게시판을 통하여 질문하고 답변하고 토론하면서 함께 얻어진 결론들과 내용들을 모아서 보고서로 만들고 그에 기반하여 중간고사와 기말고사가 제공될 것입니다.


    And we will have about 10 simple online quizzes.

    You are asked to make questions, answers, and to discuss in the BBS (QnA)  board. You will make your own report from your participations. Midterm and final exams based on those will be provided.


    특히, 매주 강의에 대응하는 내용들을 실습할 수 있는 실습실을 14개 주차별로 만들어서 실제 사용하면서 수업을 진행할 것입니다. 이 실험실들이 여러분들의 학습에 큰 도움이 될 것입니다.


    In particular, I have created 14 labs to practice the contents of each lecture.

    http://matrix.skku.ac.kr/math4ai-intro/ 


    These labs will be of great help in your learning the material.


    이와 같이 다섯 챕터에 걸친 내용을 커버하면서 고교생과 일반인들을 위한 K-MOOK 인공지능을 위한 기초수학 입문 강좌를 흥미있게 그리고 열정적으로 즐겨 주시기를 기대하겠습니다. 감사합니다.


    Please enjoy K-MOOC <Introductory Math for AI> with your own strong motivation. Thank you!





    Week 1. Introductory Mathematics for Artificial Intelligence

     

                인공지능을 위한 기초수학 입문



    인공지능에 필요한 기초수학 (1주차)

    1. 함수 그래프와 방정식의 해  14

    1.1 함수와 그래프

    1.2 다항함수

    1.3 유리함수

    1.4 삼각함수

    1.5 지수함수와 로그함수

    1.6 방정식의 해


    안녕하십니까.

    K-MOOK 인공지능을 위한 기초수학 입문을 담당하게 된 성균관대학교 자연과학대 수학과 이상구교수입니다. 반갑습니다.

    이번 고교생과 일반인을 위한 K-MOOK introductory mathematics for artificial intelligence 즉, 인공지능을 위한 기초수학 입문.


    Welcome! May I introduce myself? 


    I am Prof. Sang-Gu LEE, at Sungkyunkwan University, who is in charge of the K-MOOC <Introductory Math for Artificial Intelligence(AI)> class for everyone.



    2020 경문사에서 발간된 인공지능을 위한 기초수학 입문 introductory mathematics for AI 교재를 활용할 것이며, 그 내용은 math4ai-intro 웹 주소에서 확인할 수 있도록 제공하였고,


    This class will use the textbook <Introductory Math for Artificial Intelligence, Kyungmun Books, 2020>.

    And my lecture note can be found in the following web address:

                    http://matrix.skku.ac.kr/math4ai-intro/  .


    목차는 Part 1. 인공지능에 필요한 기초수학, Part 2. 인공지능과 행렬, Part 3 인공지능과 최적해, Part 4, 인공지능과 통계, Part5. 주성분 분석과 인공신경망에 대해서 다루고 이 전체 내용을 14주에 걸쳐서 공부하겠습니다.


    The contents of the class are given as follows:

     

    Part 1. Fundamental Mathematics for AI

    Part 2.  AI and Matrix

    Part 3  AI and Optimization.

    Part 4,  AI and Statistics

    Part 5. Principal Component Analysis(PCA) and Artificial Neural Network(ANN)


    We will study these 5 parts in 14 weeks.



    평가방법은 출석, 참여, 10회에 걸친 퀴즈에서 40%, 중간고사 30% 그리고 보고서와 기말시험 30%, 총 100% 비율로 평가할 것입니다.


    The evaluation will be divided into 40% for attendance/participation/quizzes (10 times), 30% for Midterms, 30% for Reports and Final exam, for a total of 100%.


    인공지능을 위한 기초수학 입문 내용을 좀 더 자세히 설명 드리면, 첫주차에는 인공지능에 필요한 기초수학으로 다양한 함수의 그래프를 그리고 그것을 이용하여 방정식의 해를 구하는 방법을 복습합니다. 그 다음 인공지능과 행렬에서는, 5주에 걸쳐서 선형대수학 지식을 정리한 데이터와 행렬, 데이터의 분류, 선형연립방정식의 해집합, 정사영과 최소제곱문제, 행렬분해 특히 특잇값 분해를 다룹니다. 인공지능과 최적해에서는 미분, 도함수, 편도함수 내용을 3주에 걸쳐 다루면서 극대와 극소, 경사하강법, Gradient Descent method를 이용하여 최소제곱문제의 해를 구하는 방법을 이론 및 실제 실습을 통하여 학습할 것입니다. 그리고 2주에 걸쳐서 통계의 기초 지식인 counting method 순열, 조합, 확률, 확률변수, 확률분포 그리고 기댓값, 분산, 공분산, 공분산 행렬에 대해서 학습할 것입니다. 마지막으로 주성분 분석 PCA와 인공신경망에서는 Principla Component Analysis, Artificial Neural Network 그리고 MNIST 데이터 set을 이용한 숫자인식 이론과 실습을 하고 본 강좌를 마칠 계획입니다.


    Let's talk more about  <Introductory Math for Artificial Intelligence(AI)>.


    In the first week, we will review some mathematical fundamentals that are needed in the study of AI; For example, Plotting graphs of functions and Finding the solutions of equations.


    For the next five weeks in Part 2 <AI and Matrix>, we will study basic linear algebra; Data and matrices, Classification of data, System of linear equations, Projections, the Least squares problem, Matrix decompositions including the Singular Value Decomposition.


    For the next three weeks in Part 3 <AI and Optimization>, we will study basic Calculus; Derivative, Partial Derivative, Gradient, Local Extremum, Absolute minimum, Gradient Descent method. We will use the Gradient Descent method to find the solution of a least square problem.


    For the two weeks in Part 4 <AI and Statistics>, we will study basic statistics: permutation, combination, probability, probability variable, probability distribution, and expected values, variance, covariance, covariance matrix.


    Finally, we will use two weeks for Part 5 <PCA and ANN>, you will learn Principle Component Analysis, Artificial Neural Network and an ANN example using MNIST Data set.


    본 강좌를 효과적으로 운영하기 위하여 웹 컨텐츠를 제공하면서 동시에 사이버랩, matrix.skku.ac.kr/KOFAC 한국과학창의재단과 협조하여 만든 사이버랩을 이용하여 파이센 기반의 코드를 사용하면서 실제 계산과정을 획기적으로 줄여서 데이터들을 다루고 시뮬레이션 할 수 있도록 준비하였습니다.


    We will provide Web content and CyberLab for you to use to learn the content more effectively. http://matrix.skku.ac.kr/math4ai-intro/ and

    http://matrix.skku.ac.kr/KOFAC/ 


    You will be more confident for computing and simulating your data with the provided Python based Sage codes and R codes.


    그리고 본 강좌에서는 10회에 걸쳐 간단한 온라인 퀴즈를 제공할 것이고, 질의 게시판을 통하여 질문하고 답변하고 토론하면서 함께 얻어진 결론들과 내용들을 모아서 보고서로 만들고 그에 기반하여 중간고사와 기말고사가 제공될 것입니다.


    And we will have about 10 simple online quizzes.

    You are asked to make questions, answers, and to discuss in the BBS (QnA)  board. You will make your own report from your participations. Midterm and final exams based on those will be provided.


    특히, 매주 강의에 대응하는 내용들을 실습할 수 있는 실습실을 14개 주차별로 만들어서 실제 사용하면서 수업을 진행할 것입니다. 이 실험실들이 여러분들의 학습에 큰 도움이 될 것입니다.


    In particular, I have created 14 labs to practice the contents of each lecture.

    http://matrix.skku.ac.kr/math4ai-intro/ 


    These labs will be of great help in your learning the material.


    이와 같이 다섯 챕터에 걸친 내용을 커버하면서 고교생과 일반인들을 위한 K-MOOK 인공지능을 위한 기초수학 입문 강좌를 흥미있게 그리고 열정적으로 즐겨 주시기를 기대하겠습니다. 감사합니다.


    Please enjoy K-MOOC <Introductory Math for AI> with your own strong motivation. Thank you!


    [1-2 pre]


    1주차 1차시 강의에서는 함수의 그래프를 그리는 학습을 하였습니다. 이번 2차시에서는 방정식의 해를 구하기 위해서 의 그래프를 그리거나 를 만족하는 를 찾기 위해서 그래프를 그리고 그것이 축과 만나는 해를 공학적 도구를 이용하여 그래프를 그리고 그 점을 점점 그래프를 확대해가면서 해가 어디로 수렴해 가는지 확인하는 방법을 배웁니다.

    이 방법으로 여러분이 어떤 방정식 을 만나더라도 근의 공식 없이도 방정식의 해를 구할 수 있는 자신감을 갖도록 할 것입니다.



    [1-2 pre]


    In the previous session, we learned to draw graphs of given functions. In this second session, we will learn how to find the solution of equations. At first, we draw a graph to find the solution that meets the x-axis and extend this concept. And then we learn how to find approximate solutions.


    After this, we will be confident to solve any type of equations without much difficulty.


    [1주차 2강]


    반갑습니다. 1주차 1차시 강의에서는 함수를 복습하고 함수의 그래프를 그리는 방법을 배웠습니다. 따라서 여러분은 그 동안 초-중-고등학교에서 배운 거의 모든 함수의 그래프를 간단한 명령어를 이용해서 그릴 수 있게 되었습니다. 이제 그것을 이용해서 다양한 방정식의 해(solution)을 구하는 방법을 학습하도록 하겠습니다.


    [Week 1 - Lecture 2]  Nice to meet you. In the previous session, we have reviewed functions and learned how to draw the graphs of functions. So we can draw graphs of most of the functions we have  learned (in elementary, middle, and high school years) by using a simple code. Now let's learn how to solve various equations using this technique.



    Section 1. 6  <Solutions of Equations>   1장 6절 <방정식의 해>


    방정식이란 변수를 포함하는 등식에서 변수의 값에 따라 참 또는 거짓이 되는 식입니다. 고등학교에서 다루는 방정식은 일차방정식, 이차방정식, 삼차방정식 등의 다항방정식과 유리방정식, 무리방정식, 삼각방정식, 지수방정식, 로그방정식 등이 있습니다. 이런 방정식의 해(solution)를 구하는 것은 인공지능은 물론 모든 수학에서 가장 중요한 문제 중 하나입니다. 우리는 앞에서 배운 코드를 활용하여 그래프를 그리고 근사해를 구하는 방법을 간단히 복습하도록 하겠습니다.


    An equation is a statement that asserts the equality of two expressions, so it becomes true or false, depending on the value of the variable. Equations covered in high school include polynomial equations (such as linear, quadratic, and cubic equations), rational equations, irrational equations, trigonometric equations, exponential equations, and logarithmic equations.


    Solving these equations is one of the most important problems in mathematics as well as in AI study. We will briefly review how to draw a graph and find an approximate solution using the code we have learned earlier.


    먼저 방정식의 해를 구해볼까요.


    Let's solve an equation first.


    다음과 같은 이차방정식이 주어지면, 인수분해를 해서 ‘해는 또는 ’ 라는 것을 구하는 방법을 배웠습니다. 다항식이 4차식이면 마찬가지로 인수분해를 하거나 그래프를 그려서 축과 만나는 값으로 그릴 수가 있습니다. 또 제곱근(square root, 루트)이 포함된 함수, 삼각함수를 포함하는 삼각방정식, 로그함수를 포함하는 로그방정식 등으로 점점 복잡해지는데. 이런 다양한 방정식의 해를 구하는 일반적인 방법을 알려드리도록 하겠습니다.


    For any given quadratic equation, we learned how to factor it. In the following example, the solution is or . If we solve a higher degree polynomial equation, we may also factor it or plot it and find points at which the graph meets the -axis. Let me show you a general way to solve various equations, such as trigonometric equations, logarithmic equations, irrational equations, etc.


    우리가 앞에서 소개한 python 기반의 Sage명령어 중에 solve라는 명령어가 있습니다. 우선 이 간단한 명령어 하나로 해를 바로 구할 수가 있습니다. 이때 컴퓨터가 이해하도록 등호는 두 개(==)를 사용해서 입력하는데, 만일 solve 명령어가 입력한 방정식의 해를 바로 구해주면 큰 문제가 없지만, 이 solve 명령어가 원하는 답을 한 번에 정확하게 구해주기 못하는 방정식도 많이 있을 수 있습니다. 그럴 때, 우리가 이전에 배운 수학적 지식을 이용하여 동치인 다른 방정식으로 바꾸거나 인수분해를 통해 더 쉬운 방정식으로 문제를 바꿔서 푸는 요령(skill)이 필요할 수도 있습니다.


    First, use the python-based Sage command ‘solve’ to solve a equation. We can get the solution right away with this simple command. In ‘solve’ command, two equal signs (==) must be used. If this 'solve code' immediately solves the given equation, there is no big problem at all, but there may be many other equations where this ‘solve’ cannot accurately give the answer to us. In that case, we can use our mathematical knowledge to convert this equation to a new equivalent equation.


    예를 들어서, =0이라는 방정식이면, 이 식을 주고 에 대해서 solve 명령어를 사용하여, 해를 구하라고 하면 컴퓨터는 바로 x=1, x=2라는 해를 제공합니다. 즉, 우리가 인수분해를 해서 손으로 구한 것과 같은 값이 나오죠.


    For example, if we have a given equation , then any computer immediately gives the solution x=1 and x=2 when we ask to find solutions  by the above 'solve' command. We can see that this solution is exactly same as the solution that we find by hand.


    같은 방식으로 좀 더 복잡한 4차식의 다항방정식을 주고 solve 명령어를 사용하여 해를 구해달라고 하면, 이렇게 4개의 해를 구해줍니다. 여기서 이 는 복소수입니다. 복소수 해까지도 구해줍니다. 


    In the same way, if we try to solve the following quartic equation, then we will obtain a solution by using ‘solve’ command. In this example, we have the four solutions. Here means the imaginary unit. Even complex solutions can be found with this code/command.


    그런데 이라는 식의 근을 구하려고 하면, 우리가 배운 지식을 활용하여 식을 더 단순하게 바꾼 후 명령어를 사용하는 것이 현명합니다. 즉 를 왼쪽으로 옮긴 후, 양변에 제곱을 해주면 이 되겠죠. 이 방정식을 에 대해서 풀라고(solve) 하면 x=4 와 -1로 해가 나옵니다. 이렇게 하면 앞에 쓴 명령어를 그냥 써서도 답을 쉽게 구할 수 있습니다.


    When we want to solve the equation x+ sqrt {x+5} =1 , it is better to convert this equation into a simpler form by using mathematical knowledge and then use ‘solve’ command. For example, if we move to the right and square both sides, it becomes . Then we solve it to obtain solutions x=4 and –1. In this way, you can easily get an answer by simply using the 'solve' command.


    단지, 이렇게 구한 다음에, 앞의 에 루트가 있었으니까 루트 안의 값을 언제나 0보다 크거나 같아야 되니까, 이고 solve 명령어가 준 4와 -1이 다 맞는 해가 되네요. (그런데 만약에 여기 루트 안의 값이 0보다 작은 값이 나오는 값을 컴퓨터가 준다면, 그 값은 우리가 구하는 해에서 제거하는 그런 sense도 필요합니다. )

    However, after finding candidates to be solutions by your computer code, we have to check which one is the (correct) solution. In the previous example, there was a square root of , which menas the value inside the square root must always be greater than or equal to 0, so it must be , so our candidates 4 and -1 which were found by 'Code' are both correct solutions and they forms the solution set. (By the way, if your computer gives a value where the value in the square root is less than 0, the value needs to be removed from the solution set that we are looking for.)


    또, 이런 삼각방정식의 경우, 이렇게 쉽게 인수분해를 할 수 있습니다. 인수분해 한 다음에 , 이 해가 되는데, 이때, 그래프를 그려서 가 , 가 를 구할 수도 있지만, 이런 경우는 식을 바로 집어 놓고 solve를 해서 값을 구할 수도 있습니다. 중요한 것은 그냥 solve 명령어를 이용했을 경우에는, 특수각 범위에서, 즉 와 사이에서의 답만을 준다는 것을 확인할 수 있습니다.


    In the case of this trigonometric equation, we can easily factorize this. After factoring, we have equations and . We can draw a graph of this equation to find which are and . And we can use the code ‘solve’ to find the same solutions. The important thing is that if we just use the text 'solve' in the code, we will only have particular solutions on a special angle range,  (, ).


     여기서 설명하려 했던 것은, 우리가 이미 간단한 식은 인수분해를 할 수 있으니까, 인수분해를 하고 각각의 답을 구하여 취합하는 것이 훨씬 쉬울 수도 있고 정확할 수도 있다는 것입니다.


    What I'm trying to explain here is that, since we can factor it in simple equations, it can be much easier or more accurate to factor and get the answers for each and put them together.


    그러나 코드를 이용하여 한 번에 답을 구할 수도 있습니다. 다시 확인해볼까요. 인 경우에, 우리가 앞에서는 특수각만 구했으니까 이번에 일반각을 구하려면, 옵션으로 to_poly_solve = 'force' 라는 명령어를 활용할 수 있는데,


    However, there is a code where we can have all general solutions (the solution set) at once. Let’s check it again. In the case of , we only found the particular solution before, so we will try to find all the general solutions at this time. It can be done by adding <to_poly_solve ='force'> as an option in a new code 'solve'.


    이 명령어를 이용해서 일 때 일반해를 구하라고 교재에서 제시한 명령어 solve(sin(x) == -1, x, to_poly_solve = 'force') 를 주어서 solve 하라고 하면, 다음과 같이  이라는 일반해를 구해줍니다. 이 공학적 도구가 주는 답은 z2, z5, z6 인데 이는 주기를 나타내는 을 변수로 제공해 준 것입니다. 또 의 일반해를 구하라고 명령어 solve(sin(x) == 1/2, x, to_poly_solve = 'force') 를 주면, 다음과 같은 해가 나오는데, 그 의미가 가 라는 의미입니다. 즉 두 개의 주기해 (periodic solution)을 구해줍니다.


    If we use the code <solve(sin(x) == -1, x, to_poly_solve ='force')> in http://matrix.skku.ac.kr/KOFAC/ when , it gives the general solution , where is integer. In the answer obtained from this code, z2, z5, z6 mean any integers. Also, if we use the code <solve(sin(x) == 1/2, x, to_poly_solve ='force')> when , it gives the general solutions which are two periodic solutions of .


    다음 를 일단 단순하게 정리를 해보죠. 이 되고, 로그는 덧셈이 곱셈으로 표시되니까 . 즉, 이 된다는 것을 알 수 있고, 그럼 간단한 식인 이것만 풀면 답이 구해집니다. 그리고 와 가 0보다 크다는 조건 하에서 이라는 그래프를  그려서 해를 구해보면, 우선 다음과 같이 show(y) 명령어로 다음과 같은 그래프가 그려집니다. 그렇다면 축과 만나는 여기와 여기에서 해가 존재한다는 의미이고, 자세히 보면 -4 와 2 가 되는 것을 확인할 수 있죠. 즉, 이 그래프를 그려보면 -4와 2가 의 해의 후보가 됩니다. 그렇다면 원래 우리가 찾고자 하는 로그 방정식의 해는 –4와 2 중에 있는데, 와 가 항상 0보다 커야 되니까, -4는 실제 답에서 배제되고 2만이 해가 되는 것입니다.


    Next, let's simplify . It becomes , and because of the property of logarithm, this can be represented as . That means , and then the answer can be easily obtained. If we draw a graph of under the condition that and are greater than 0, then we will see the solution. First, the following graph is drawn with the code <show(y)>. Then the solution exists where the graph meets the -axis, and you can see that it becomes -4 and 2. In other words, if you plot this graph, -4 and 2 become candidates for the solution of . Then, the original solution of the logarithmic equation we are looking for is -4 or 2, but since and must always be greater than 0, -4 should be excluded to make the solution set is having only {2}.


    즉, 지수함수, 로그함수, 유리함수 등 다양한 방정식이 있어도 단순하게 고쳐서 (파이썬 코드를 잘 몰라도) 간단한 solve 명령어로 해를 구한 후, 우리가 배운 수학 지식을 활용하면 근을 쉽게 구할 수 있다는 의미입니다.

     

    In other words, even for various equations such as exponential equations, logarithmic equations, rational equations, etc., we can (1) use simply those equations and (2) get all candidates to be solutions with a simple code 'solve' (even if you don't know Python code well), and then (3) find real  solutions from them by checking the conditions with our math knowledge.


    다음 예제는 solve 명령어로 해결하지 못하는 방정식의 경우는, 그래프를 그려서 근삿값을 구하는 방법으로 우리가 원하는 해를 구할 수 있다는 것을 보여준 것입니다.

    The following example shows that for equations that cannot be easily solved by hand or with simple codes, we can still find an approximate solution by drawing a graph and then use approximation skill in the graph.


    예를 들어서 다음과 같이 복잡한 방정식이 있다고 합시다. 이런 방정식은 고등학교까지 배운 수학지식으로는 풀기가 쉽지 않습니다. 명령어를 통해서 일단 solve 해서 구하면 답이 이렇게 두 개가 나타납니다.

    For example, assume that we have an equation . This equation is not easy to solve by hand. So once we solve it with a 'code'. Then we will have the output (2 strange expressions) like the following.


    이게 의미하는 바는 가 이런 식과 이런 식으로 나타난다는 의미인데. 실제로 우리가 원하는 것은 이런 복잡한 수식이 아니라, 대개 어디쯤에 이 값이 존재하는지가 필요합니다. 기존의 solve 명령어는 단순히 이런 모양의 해를 주고, 많은 경우는 해를 주지 못하기도 합니다. 이 때 아래와 같은 방법이 효과적입니다.


    What we really want is not such a form, but where this value(solution) exists. The existing code 'solve' often simply gives strange type of expressions, furthermore solution does not even exist in many cases. In that case, the following method will be very effective.


    그래프를 그려서 확대하면서 더 쉽게, 더 정확한 근사해를 구할 수 있습니다.

    We can very easily get more accurate approximate solutions by enlarging a graph.


    위의 식의 경우, 항을 모두 왼쪽으로 옮겨서 다음과 같이 쓸 수 있습니다. 그러면 우리는 언제나 로그함수, 지수함수, 삼각함수를 다 그릴 수 있으니까 의 그래프를 그릴 수 있습니다. 이것이 축과 만나는 값을 찾으면 되지요.


    In the case of the above equation, we can write by shifting all expressions to the left as follows. Then we can always plot all logarithmic, exponential, and trigonometric functions, so we can draw the graph of . We just need to find the -value where this graph meets this -axis.


    이를 위하여 적당한 범위에서 함수의 그래프를 확대해 보면 다음과 같습니다. 즉,  를 -2하고 2 구간에서 그려보면, 다음과 같이 이 값을 지나는 걸 확인할 수 있습니다.


    For this, we enlarge the graph of the function within the proper range. For example, if you draw the graph of at the interval of -2 and 2, you can see that it passes this value as follows.


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    plot(4*exp(-x^2)*sin(x) - x^2 + x - 1, (x, -2, 2))

    ----------------------------------------------------------------------

    그림입니다.
원본 그림의 이름: mem00004e440009.tmp
원본 그림의 크기: 가로 630pixel, 세로 470pixel     ■

                                                             <---화면에 있으니 자막에는 없어도 됨)




     그러면 이 값이 몇쯤 되나 자세히 보면. 하나의 해(solution)는 0.1과 0.3 사이에 있다는 걸 대충 알 수 있으며, 또 두 번째 해(solution)는 1과 1.5 사이에 있는 것을 알 수 있죠. 즉, 두 개의 근이 있는데, 두 개의 근이 이 사이에 있다는 걸 알았으니까, 예를 들어서, 이 사이에 있는 근을 더 정확하게 파악하기 위해서, 이 구간을 0.2에서 0.23 정도로 가깝게 잡아서 그려줍니다. 즉, 확대한다는 의미입니다. 확대해서 그림을 그려주면 여기를 지난다는, 여기가 첫 번째 근이라는 걸, 해인 걸 알 수 있죠. 그럼 대략 0.21678쯤 된다고 볼 수 있습니다. 즉, 0.219 근처에 해가 있다는 것을 알 수가 있습니다.


    Then, looking closely to see where this value is located, we know that one solution is in between 0.1 and 0.3, and the second solution is in between 1 and 1.5. In other words, there are two roots where are in between these two interval. In order to more accurately determine the roots between them, draw the graph with the small interval (0.2, 0.23). We will see a bigger graph (like we see it with a microscope, Magnifying). Then we can see the graph passes through here, where it is the first solution(root). It can be said that it is near 0.22. Finally we can see that there is a root around 0.219.


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    plot(4*exp(-x^2)*sin(x) - x^2 + x - 1, (x, 0.2, 0.23))

    ----------------------------------------------------------------------

    그림입니다.
원본 그림의 이름: plot-2.jpg
원본 그림의 크기: 가로 629pixel, 세로 470pixel  (즉, 0.219 근처에 하나의 해가 있음을 알 수 있다.)  ■


    이런 식으로 plot의 구간을 조정하여 점점 확대하면서도 근을 구할 수 있으니까, 구하는 해에 점점 더 가까운 근사해를 구할 수 있습니다.


    In this way, we can find all roots that we need by gradually adjusting the interval of our plot, so we can get an approximate solution as close as we want.


    이는 우리가 그림만 그릴 수 있으면 그리고 기본적인 그동안 우리가 배운 수학적 지식을 활용해서 주어진 식을 좀 더 단순화 시켜서 그림만 그리고, 확대해가면, 우리가 배운 모든 함수들을 포함하는 방정식의 충분히 정확한 근사해를 구할 수 있는 것입니다.


    It means that if we can plot the graph of a given equation (and Magnifying) after simplifying with our basic mathematical knowledge, then we can obtain a sufficiently accurate approximate solution of the given equation that includes all the functions we have learned.


    이 식에 find root를 하라고 새로운 명령어를 주면, 근도 이렇게 구해질 수 있지만, 여기서 구한 값이 0.219163같이 우리가 그래프를 그려서 얻은 값과 같은 값임을 알 수 있습니다.


    If we give a new Code <find_roots>, the root can also be obtained like this, but we will see that wether the value obtained here is the same as the value obtained by plotting the graph, such as 0.219163.


    이런 식으로 다양한 명령어를 이용해서 답을 구할 수 있고, 여러분들이 그림만 그려서도 언제나 근을 구할 수 있습니다. 그걸 점점 확대해가고 한다는 것이 단계를 거쳐서 근을 근사식을 구하는 과정과 일치하고, 사실 위의 과정 포함하는 알고리즘을 모두 입력한 것이 find_root라는 명령어입니다.


    By this way, we can see there are a variety of codes that helps to get solutions of equations, but we know that we can always find their solutions by just plotting a graph and magnifying it. The Code <find_root> was made to execute the process explained above.


    이 절에서 배운 것은 앞에서 배운 그래프 그리는 것을 이용하여, 우리가 그동안 배운 함수가 포함된 다양한 다항식과 삼각방정식, 유리식 등 어떤 경우도 그 함수를 적절히 단순화시켜서 그래프를 그려서 그것을 확대 축소하면서 근을 구할 수 있다는 것을 확인한 것입니다.


    What we learned in this section is that we can plot a graph of any functions we have learned so far, and obtain all the existing roots by appropriately magnifying the graph.


     find_root 명령어는 앞에 우리가 해본 전체 과정을 하나로 묶은 명령어입니다.


    The <find_root> code is what combines the entire process we did earlier.


    결론적으로, 이렇게 그래프를 그리고 그래프를 축소, 확대해서 근사해를, 대개 0.22와 1.08임을 대충 알 수 있었죠. 이런 값이 보통 우리가 원하는 해입니다.


    In the above example, we can plot a graph and zoom in or out to get an approximate solution of the equation, roughly 0.22 and 1.08. This is the solution we want.


    오늘 살펴본 내용을 활용하면 여러분이 이제 초-중-고등학교 그리고 사실 대학에서 배우는 내용까지 포함하여, 대부분의 함수 그래프를 그릴 수 있으며, 그 동안 배운 모든 대부분의 방정식에 대한 해, 근사해를 쉽게 구할 수가 있습니다.

     

     By using what we learned today, we can now plot the graphs for the most of the functions, and we can easily find solutions or approximate solutions for most of the equations we need to solve. This is the key feature of this lecture.


    여기 주소 http://matrix.skku.ac.kr/2020-Math4AI-PBL/ 을 보시면, 이런 방식으로 배우고, 다양한 문제를 푸는 시도를 한 학생들의 기록들이 제시되어 있습니다. 참고로 보시기 바랍니다. 그리고 주위에 있는 다양한 방정식, 무엇이든 가져오셔서 그 방정식의 그래프를 그리고 그 그래프의 근이 어디 있는지 대충 확인한 후에, 그것을 확대해가면서 구간을 해에 가깝게 점점 줄여주면서, 즉 그래프를 확대하면서 근사해를 찾는 경험을 해보시면 앞으로 배울 [인공지능을 위한 기초수학 입문] 학습에 큰 도움이 될 것입니다.

    If you take a look at the following link

    http://matrix.skku.ac.kr/2020-Math4AI-PBL/, there are many records of students who tried to solve various problems. Please refer to this for references. It will be a big help for your learning of [Introductory Mathematics for Artificial Intelligence].


    첫 번째 주 과제로는 위의 <열린 문제>를 풀어본 후 <간단한 자기소개와 수강 동기> 및 본인이 요약, 실습, 질문, 답변한 내용을 문의게시판에 공유하시는 것으로 하겠습니다. 이미 교재 http://matrix.skku.ac.kr/math4ai-intro/W1/ 를 미리 공개해 주어서, 수업 전에 자기소개와 자기가 해본 것을 공유한 학생들도 있을 수 있을 텐데. 그것이 바로 여러분들에게 제가 기대했던 내용입니다.


    The first week assignment is to solve the above <Open Problem> and then share <Self-Introduction and your Motivation to take this course>.

    Hope you can share it in the QnA board of LMS.


    Practice the contents at http://matrix.skku.ac.kr/math4ai-intro/W1/. That's what I expect to you to do.


     아래 사진은 수학분야 필즈상(Field) 수상자들 중에서 제가 스탠포드 대학을 방문했을 때, 그 대학의 교수 중 수학의 노벨상이라고 불리는 필즈상을 받은 분들 사진입니다. 그에 보태서 코로나-19 때문에 위축되었던 우리가 더 힘내자는 사진이었습니다.


    The .\PICture below shows Professors at Stanford University who received the Fields Medal (the Nobel Prize in Mathematics). I took a photo when I visited that University in January. The other one was a .\PICture that we wanted to cheer up students who try to overcome this COVID-19 situation.


    이제 오늘 배운 내용을 복습해보도록 합시다.


    Now let's review what we have learned today.


    1주차 <함수 그래프와 방정식의 해>에서는 <사이버실습실>을 활용해서 그래프를 그리고 방정식의 해를 구하는 과정을 실습했습니다.


    In the first lecture, <Graphs of Functions and Solution of Equations>, we practiced the process of drawing the graphs and solving equations using the <Cyber ​​Lab>.


     순서는 함수의 개념을 소개한 후, 함수의 그래프를 어떻게 그리는지를 학습했고. 그래프를 그리는 skill을 이용하여 이런 함수들이 포함된  방정식의 해(solution)을 그래프를 통해서 대략 어디 있는지 파악하고, 그것을 확대해가면서 점점 그 해가 어디로 수렴하는지를 확인하는 과정을 실습실에서 학습하였습니다.

    구체적으로는 다항함수, 유리함수, 삼각함수, 지수함수, 로그함수를 소개하고, 또 그래프를 어떻게 그리는지를 학습했으며, 그래프 그리는 것을 이용해서 이런 함수들이 포함(involve)된 다양한 다항식의 해를 명령어로 구할 수도 있지만, 명령어가 쉽게 구해질 수 없는 경우에도 그래프를 그리고 확대하면서 답을 구할 수 있다는 것을 확인했습니다. 다음과 같은 그래프들이 우리가 학습한 내용이었습니다.


    We have introduced to the concept of a function, and learned how to plot the graph of any given one variable function. We used this skill to find the solution set of equations, such as polynomial equations, rational equations, trigonometric equations, exponential equations, and logarithmic equations. The solution can be found not only by using 'Code', but also by plotting a graph and zooming in, even though we don't know many computer language. The following graphs were what we learned.


     다음 시간에는 2주차(week 2) <데이터와 행렬>을 시작하도록 하겠습니다.


    In the next class, we will study <Data and Matrix>.




    그림입니다.
원본 그림의 이름: KakaoTalk_20200619_175748331.jpg
원본 그림의 크기: 가로 5304pixel, 세로 1540pixel

     

     

     

    Week 2. Introductory Mathematics for Artificial Intelligence

     

                인공지능을 위한 기초수학 입문


    II.  인공지능과 행렬 (2-6주차)        25

    2.  데이터와 행렬 25

    2.1 순서쌍과 벡터

    2.2 벡터 연산

    2.3 행렬과 텐서

    2.4 행렬 연산

    2.5 행렬의 연산법칙


    [2-1 pre]


    안녕하십니까. [인공지능을 위한 기초수학 입문] 2주차, <데이터와 행렬>을 학습하도록 하겠습니다. 이번 주에는 데이터를 표현하는 유용한 도구인 벡터와 행렬에 대하여 학습할 것입니다.


    Hi!! In this Week 2 on <Data and Matrix>, we will study vectors and matrices, which are useful tools for representing data.


    하나의 데이터는 벡터로, 여러 데이터는 행렬로 나타낼 수 있습니다. 또한 이미지 데이터도 행렬과 이를 확장한 개념인 텐서(Tensor)로 표현할 수 있습니다. 우리는 벡터와 행렬의 기본 개념과 성질 및 연산을 익혀 이를 이용하여 데이터를 표현하는 방법에 대하여 학습하게 됩니다. 이는 앞으로 전개될 인공지능 수학에 기초가 됩니다.


    One data can be represented as a vector, and a number of data can be represented as a matrix. Also, images can be expressed as matrices and tensors, a concept that extends matrices. We learn some basic concepts, properties, and operations of vectors and matrices, and learn how to represent data using them. This will be the basis for upcoming chapters.


    이번 주에는 <데이터와 행렬>이라는 제목으로 1절에서 순서쌍과 벡터, 2절에서 벡터 연산, 3절에서 행렬과 텐서, 4절에서 행렬 연산 그리고 마지막으로 행렬의 연산 법칙을 학습하게 됩니다.


    We will learn about ordered pairs and vectors in Section 1. vector operations in Section 2, matrices and tensors in Section 3, and matrix operations and their properties in Section 4.


    주어진 데이터로부터 그 데이터를 벡터 또는 행렬로 표현하고 그 행렬들 사이에 덧셈, 뺄셈, 곱셈 연산을 정의한 후에 전치 행렬, 역행렬, 대각선 행렬 등 다양한 행렬들을 소개하고, 그 지식을 이용하여, 어떻게 역행렬을 찾고, 어떻게 전치 행렬을 찾고 하는 다양한 행렬 연산 문제를 다루게 될 것입니다.


    We learn how to express the given data as a vector or matrix, how to define addition, subtraction, and multiplication between the matrices, and introduce various matrices such as transpose matrices, inverse matrices, and diagonal matrices, and find the inverse matrices. We deal with various matrix operation problems.


    [2주차 1강] [Week 2, Lesson 1]

    반갑습니다. K-MOOC, 인공지능을 위한 기초수학 입문 2주차 강의입니다.


    [Week 2, Lesson 1] Nice to see you in the K-MOOC introductory mathematics for AI class.


    1주차에서는 함수의 그래프를 그리는 것과, 함수의 그래프를 그리는 것을 이용해서 방정식의 해를 구하는 방법에 대해서 학습을 했습니다.


    In the previous lecture, we learned how to plot the graph of a function and how to solve an equation from the graph.


    이번 2주차부터는 그야말로 인공지능 수학의 시작인 <인공지능과 행렬>에 대해서 학습하도록 하겠습니다.


    From this week, we will study <AI and Matrices>. This will be the starting point to our journey of the <Mathematics for AI>.


    이 행렬부분을 다루는 학문을 선형대수학이라고 하는데, 선형대수학은 우리가 배우는 수학 중 가장 유용한 수학이라고 알려져 있습니다. 컴퓨터를 활용하는 거의 모든 응용수학은 전체 또는 부분적으로 행렬계산에 의존합니다. 이 장에서는 인공지능에 필요한 선형대수학의 기본내용과 데이터 차원 축소에 꼭 필요한 특잇값 분해(SVD)에 대하여 학습합니다.


    Linear algebra is the branch of mathematics concerning matrices, which is known to be the most useful. Almost applied mathematics with computing relies on matrix computations. In this part of <AI and Matrices>, we will learn the fundamentals of linear algebra required for artificial intelligence, including singular value decomposition (SVD), which is essential for dimension reduction of data.


    2주차 강의 내용의 실습실은 2주차(week2, W2) 실습실에 준비되어 있으며, 오늘 강의를 진행하면서 실습실에서 실제로 행렬연산에 대한 실습을 해 보도록 하겠습니다. http://matrix.skku.ac.kr/math4ai-intro/W2/


    In the cyber lab, the contents of this lecture is prepared at the link http://matrix.skku.ac.kr/math4ai-intro/W2/. In this lab, we will practice the matrix operations.


    2주차, 데이터와 행렬. 순서쌍과 벡터부터 시작하겠습니다.


    Let's start with Section 1, <ordered pairs and vectors>.


    데이터는 순서쌍(ordered pair, 순서조 -tuple)으로 표현할 수 있습니다. 예를 들어, 어떤 사람의 키, 몸무게, 연령, 성별 등은 그 사람에 관한 데이터가 될 수 있습니다.  이 데이터를 다음과 같이 순서쌍(순서조, -tuple)으로 나타낼 수 있습니다. 여기서 키, 몸무게, 연령, 성별 각각을 데이터를 이루는 성분이라고 합니다.


    Data can be represented as an ordered pair (or n-tuple). For example, the information on height, weight, age, sex of a person can be considered as a data for that person. This data can be expressed as an ordered 4-tuple. Each number is called a component of the data.


      그래서 김모씨의 키는 160이고, 몸무게는 80이고 연령은 19세고, 성별은 1, 남성입니다. 박모씨는 키는 180이고, 몸무게는 56이고, 연령은 30이고 성별은 1 또는 2입니다. 하면 따라서 김모씨의 정보(인포메이션)는 다음과 같은 데이터로 표현됩니다. 이 데이터가 바로 순서쌍 또는 n-tuple, 4-tuple 즉 순서조로 표시되는 것이죠.


    We know that Kim's height is 160, his weight is 80, his age is 19, his sex is 1, which means male and Park's height is 180, his weight is 56, his age is 30, and he is a male. Therefore, the information on Mr. Kim is expressed as the following data (160, 80, 19, 1). This data could be represented in ordered 4-tuple.


    이제 이런 데이터로부터 학습을 시작하겠습니다.


    Now let's start with this data.


    성분이 2개 (또는 3개)로 이루어진 2차원 또는 3차원 데이터는 좌표평면(좌표공간)상의 한 점을 나타냅니다. 이제 4차원 이상의 데이터를 우리 눈으로 보도록 시각화할 수 없지만, 고차원 공간상에 놓인 점이라고 생각할 수 있습니다.


    Two-dimensional (or three-dimensional) data consisting of two (or three) components represents a point on the coordinate plane (or coordinate space). It is not easy to visualize a four or higher dimensional data, but we can think of it as a point lying in a higher dimensional space.


    예를 들어서, 아래에서 , 를 성분으로 하는 데이터 는 좌표평면 상의 한 점 를 나타내는데, 이때 시작점을 원점(origin) 와 끝점 로 하는 화살표로 나타낸 것을 벡터(vector)라 하고, 다음과 같이 표기합니다. 벡터를 이루는 각각의 성분은 하나의 숫자로 이루어져 있는데 이를 스칼라(scalar)라고 합니다.


    For example, the data with components and represents a point on the xy-coordinate plane. Here, an arrow whose starting point is the origin and whose end point is the point is called a vector, and is expressed as follows. Each component of a vector consists of a number, which is called a scalar.


     원점(origin)에서 점까지의 벡터를 표시할 수 있습니다. 같은 방법으로 원점에서 라는 3-tuple까지를 벡터로 표시할 수 있습니다. 공간에서 원점과 한 점 간에 표시하는 화살표(arrow)를 벡터로 표시할 수 있는 것과 마찬가지입니다.


    We can plot a vector as an arrow from the origin to point . In a similar way, we can express a vector from the origin to point . We can think this arrow in a space between the origin and a point A as a vector.


    다음 절에서 벡터 그리고 벡터들의 연산을 생각해 보겠습니다.


    In the next section, we will consider vector operations.


    벡터에는 다음과 같은 연산이 정의됩니다.


    The following operations can be defined for vectors.


    실수배(scalar multiplication), 실수 와 벡터 에 대하여 다음과 같이 곱합니다. 에 를 곱하면, 라는 벡터에 두 배를 곱하면, 이렇게 또는 로 표시할 수가 있습니다.


    (1) Scalar Multiplication:  k(a, b) = (ka, kb) for a scalar and a vector .  For example, 2 = .

     

    또 덧셈(벡터 sum 또는 vector addition)의 경우, 두 벡터를 더할 때는, 각각의 성분을 더하면 됩니다. 첫 번째 성분에 있는 와 를 더해서 첫 번째 성분에 두고, 두 번째 성분에 있는 와 를 더해서 두 번째 성분 자리에 넣으면 됩니다. 같은 식으로 연산을 할 수 있습니다.


    (2) Vector addition (vector sum): When adding two vectors, simply add each component of them. (a, b) + (c, d) = (a+c, b+d). 


    These are easy operations that you can do well.


    벡터의 연산법칙, 의 벡터 , , 와 스칼라 , 에 대해서 다음과 같은 성질들이 성립합니다.


    The following properties of vector addition and scalar multiplication hold. For any vectors , , in and any scalar , , we have


     , 이를 보통 벡터의 덧셈에 대한 교환법칙이라고 합니다.


    (1) , This is the commutative law of vector addition.


    두 번째는 결합법칙, 세 번째 경우는 ‘어떤 벡터에 영(zero) 벡터를 더하면 자기 자신인 x 가 된다’입니다.  이때 이 영(zero)을 영벡터라고 부릅니다. (영벡터)는 성분이 모두 0인 벡터를 의미합니다.


    (2) . This is the associative law.

    (3) , where (zero vector) means a vector with all zero in each components.


    또 x 에 (-x)라는 벡터를 더하면 영벡터가 됩니다. 이때 –x를 x의 음(negative) 벡터라고 표시합니다.


    (4) . If you add a vector (-x) to x, it becomes a zero vector. We call this vector -x as the negative vector of x.


    이런 식으로 선형대수학에서 기본적으로 사용하는 성질들을 소개한 것입니다.


    We introduced the basic properties of vectors that will be often used.


    결합법칙들 세 가지와 벡터 x에 상수 1을 곱하면 자기 자신이라는 스칼라 배(multiple)에 대한 항등원에 대한 성질 등입니다.


    The properties (5), (6), (7) are related with associative laws and the property (8) is related to the identity for the scalar multiplication.


    이런 성질을 활용해서 여러분들은 두 벡터 , 와 스칼라 이 있을 때, 벡터 합(addiction)와 스칼라 배를 계산하는 문제들을 배웠습니다.


    Using these properties, we computed a vector addition and scalar multiplication when two vectors , , and scalar are given in [Example 1].


    두 벡터가 있으면 성분들끼리 더해서 (0, 5, -3)을 구하면 되고 세 배해주면 원래 벡터 v에다가 3배씩 성분별로 3배씩 해서 다음 값을 구했습니다.


    For given two vectors, we added them (component-wise) to obtain (0, 5, -3), and we multiplied the original vector v=(1, 2, -4) by 3, that is, each components will be tripled.


    이것을, 2차원과 3차원 등 간단할 때는 손으로 계산하면 되지만 우리는 앞으로 빅데이터(big data)와 같이 손으로 쓰기조차도 어려운 큰 크기의 데이터를 다룰 테니까, 이를 다룰 수 있는 코드를 학습하도록 하겠습니다.

    For vectors in 2D and 3D, we can calculate them by hand, but for high dimensional vectors, it is very tedious to compute by hand. So let's learn how to do vector operations by using <Code>.


    벡터 v를 다음과 같이 (1, 2, -4)로 정의하고 벡터 w 를 (–1, 3, 1)로 정의하고 스칼라를 3으로 정의한 후, 주어진 대로 코드로 정의하고  v+w를 구해서 프린트해라, k 에다 v를 곱해서 프린트해라고 명령을 주면,

     v + w는 (0, 5, -3), kv는 (3, 6, -12)가 되어, 손으로 구한 답과 일치하는 답이 나옵니다.


    We define vectors v=(1, 2, -4), w=(–1, 3, 1) and scalar k=3. Then we write a code to compute v+w and multiply v by k, and print the result as follows. We obtain v + w = (0, 5, -3) and kv = (3, 6, -12). It is easy to see that the output matches with our answer obtained by hand computation.


    앞으로는 벡터 곱과 스칼라배를 할 때는, 언제든지 이 코드를 이용하여 여기서 벡터만 바꿔주시고, 스칼라만 바꿔주시면 같은 연산을 하실 수 있을 것입니다.


    From now on, when doing vector addition and scalar multiplication, we can always perform the same operation by only changing components of vectors and the scalars in this code.


    더구나 벡터와 스칼라는 임의로 생성하는 코드를 소개해주니까, 이 코드를 카피하셔서 여기서 실습하시면 됩니다.


    Moreover, we can randomly generate vectors like the code v = random_vector(ZZ, 7, x = -10, y = 10), so all we have to do is to copy and paste this code and practice in http://matrix.skku.ac.kr/math4ai-intro/W2/.


    -10부터 9까지의 정수를 성분으로 하는 7차원 벡터를 임의로 생성해봅시다.

    Let's create a 7D vector with components randomly chosen from -10 to 9.


    그러면 정수(ZZ, integer)의 독일 단어 첫 자가 Z입니다. 7차원 벡터로 그 안의 숫자는 –10과 9 사이에 있는 숫자들로 랜덤(random)하게 벡터를 생성해라.


    ZZ means integers since Z is the German letter for integers. The code randomly generated a 7-dim vector v with numbers between -10 and 9.


    그리고 다음 벡터는 7차 벡터를 정수에서 –5와 4 사이에서 생성해라, 그러면 두 개의 벡터가 생성됩니다.


    And the next line, w = random_vector(ZZ, 7, x = -5, y = 5) creates another 7-dim vector w whose components are integers randomly chosen from -5 to 4.


     그 다음에 스칼라 k를 –5와 4 사이에서 골라라. 그렇게 골라진 벡터 v와 w 스칼리 k를 출력해라. 그러고 나서 v + w를 벡터 합(addition)과 k 곱하기 v 스칼라 배(multiplication)을 계산해라.


    Then randomly choose a scalar k between -5 and 4 and Print the chosen vectors v, w and scalar k on the screen. Then compute the vector addition v + w and the scalar multiplication k times v.


    그러면 랜덤하게 7차원 벡터 v가 생성됩니다.

    The code will randomly generate a 7-dimensional vector v.


    그 다음에 랜덤하게 7차원 벡터 w가 생성됩니다.

    And a 7-dim vector w was randomly generated.


     그리고 스칼라가 –5와 4 사이에서 4로 생성됐죠. 그래서 두 개의 벡터를 더한 것은 다음과 같이 –6과 –3을 더했으니까 –9, -2와 4를 더했으니까 2 하는 식으로 구해졌고, 여기 v에다 4배 해준 벡터 4v는 마이너스 4, 6에 24, -2에 4는 –8, 이런 식으로 해서 구해질 수 있죠.


    And the scalar 4 was generated. Then the addition of two vectors, v+w  is as follows. Next, the scalar multiplication 4v was also given.


       이와 같이 랜덤하게 구한 벡터들 사이에 벡터 합(addition)과 스칼라 배(multiplication)을 할 수 있었습니다.


    Vector addition and scalar multiplication between randomly obtained vectors and scalars can be evaluated like this.

     

    이 벡터의 차원이 7이 아니라 20, 30, 50이었어도 가능한 것입니다.


    This can be done not only for 7D vectors, but also for higher dimensional vectors, for example, 50D vectors.


    다음 3절에서는 행렬과 텐서에 대해서 학습하도록 하겠습니다.


    In the next section, we will learn about <Matrices and Tensors>.


      앞서 언급한 어떤 사람의 키, 몸무게, 연령, 성별에 대한 데이터를 하나로 모아 다음과 같이 직사각형 모양 array로 배열할 수 있습니다. 이것을 행렬이라고 합니다.


    The data of the height, weight, age, and sex of many persons can be gathered and be arranged in a rectangular array as follows. This is called a Matrix.


    즉 첫 번째 사람의 정보(information), 두 번째 사람의 information, 세 번째 사람의 information을 줄 때 첫 번째 성별은 남자고 두 번째 여자 세 번째는 남자인 사람의 키가 160이고 170이고 180인 데이터입니다.


    The first row of this matrix is the data for the first person, and the second row for the second person etc. The first row tells us that he is the person of height 160, weight 80, age 19, and a male.


    이 데이터들을 벡터, 즉 행벡터로 하는 행렬을 생각해봤습니다. 3 by 4 행렬이 만들어졌습니다.


    So, we can create a 3 by 4 matrix whose rows are the data of each person.


     이때 행렬의 가로를 행(row) 라고 하고  세로를 열(column) 이라고 합니다.


    In this matrix, rows are "horizontal" collections of numbers and columns are "vertical" collections.

    그래서 벡터는 행벡터는 1 by n 행렬로 생각할 수 있고, 열벡터는 n by 1 행렬로 생각할 수 있습니다.


    So, a row vector can be thought of as a 1 by n matrix and a column vector as an n by 1 matrix.


     그래서 우리는 이제 ‘3 by 4 행렬을 만들었습니다’.  즉 생성했습니다.


    So, we have constructed a 3 by 4 matrix from the given data.


    이 행렬은 디지털 이미지를 유용하게 나타낼 수 있는데, 예를 들어, 디지털 이미지를 확대할 때 나타나는 작은 격자를 픽셀(pixel, 화소)이라 하는데, 각 픽셀에는 이미지의 밝은 정도를 나타내는 숫자가 들어 있다고 볼 수 있습니다. 숫자로 밝기를 생각할 수 있죠. 따라서 흑백 이미지는 하나의 행렬로 나타낼 수 있고, 컬러 이미지는 빨강(Red), 녹색(Green), 파랑(Blue), 3개의 채널(channel)로 표현된 3차원 행렬로 생각할 수 있습니다. Cube 형태의 행렬로도 생각할 수 있습니다. 이런 것을 3-D 텐서(tensor)라고 부르기도 합니다. 즉, 이런 사진 이미지, 흑백 이미지는, 밝기에 따라 숫자로 표시해서 행렬로 표시할 수 있고, 칼라 이미지라 그러면, Red, Green, Blue 세 개가 겹쳐서 만들어 내는, 다음과 같이 이렇게 칼라 이미지도 행렬들의, 3차원 행렬로 생각할 수가 있는 것이죠. 흑백이미지에 대한 정보(information), 컬러이미지에 대한 information은 다음 주소에서 찾아볼 수 있습니다.


    A matrix can be very useful for representing a digital image. For example, a small grid that appears when a digital image is enlarged is called a pixel, and each pixel contains a number indicating the brightness of the image. We can think of brightness as numbers. Therefore, a grayscale image can be represented as a matrix, and a color image can be thought of as a three-dimensional matrix represented by three channels, <Red, Green and Blue>. We can also think a color image as a cube-shaped matrix. This is called a 3-dim Tensor. In other words, such photographic images, a grayscale image can be expressed as a matrix by matching components of the matrix to the brightness of the image, and a color image can be thought of as a three-dimensional matrix consists of overlapping three components, <Red, Green and Blue>. Information on grayscale images and color images can be found at the following link. [Color image]

    https://lisaong.github.io/mldds-courseware/01_GettingStarted/numpy-tensor-slicing.slides.html


    참고로, 최근의 인공지능 기계학습 시스템은 텐서를 기본 데이터 구조로 사용합니다. 구글의 텐서플로우(TensorFlow)가 잘 알려져 있는데, 이름도 텐서에서 따온 것입니다. 그렇다면 텐서(Tensor)란 무엇일까? 텐서를 간단히 말씀드리면, 텐서는 데이터를 한군데 모아 놓을 수 있는 저장소인 컨테이너(container)라고 말할 수 있고, 대부분 수치형 데이터를 다루므로 숫자를 위한 컨테이너라고 이해할 수 있습니다. 1-D 텐서가 벡터이고, 2-D 텐서가 행렬이고, 행렬의 일반화된 모습을 텐서라고 생각하시면 되겠습니다.


    For reference, a machine learning systems are using Tensors as their basic data structure. The well known Google's TensorFlow takes its name from this term. So what is a Tensor? In short, it can be said to be a container, which is a storage that can put data together, and since most of them deal with numeric data, it can be understood as a container for numbers. We can think of a 1-dim Tensor as a Vector, a 2-dim Tensor as a Matrix, and the generalized form of a Matrix as a Tensor.


    행렬 연산, 간단한 행렬연산을 소개하도록 하죠. 행렬에는 다음과 같이 세 가지 연산이 정의됩니다. 먼저 스칼라배(scalar multiplication), 실수배. k배, 2배. 어떤 행렬에 2배 해준다고 그러면 성분(entry)에 각각 두 배씩 해주면 되는 거구요.


    [Matrix Operations] Let's introduce simple matrix operations. Three operations are defined for matrices: first, scalar multiplication(real multiplication), k times(2 times). If you multiply a certain matrix by a scalar 2, then you just have to double each entry in the matrix.


    행렬 합(addition) 즉 행렬 사이의 덧셈을 생각하며, 먼저 행렬의 크기가 같이 n by m 행렬이어야 하니까, 같은 사이즈 행렬 두 개를 더 한다 그러면 각 성분 각각을 더해주면 됩니다. 벡터 합(addition) 하고 똑같죠. 행렬 두 개가 더해지면 그대로 성분끼리만 더해지면 됩니다. 다음이 행렬의 곱셈 곱(product)입니다.


    Matrix addition is the operation of adding two matrices by adding the corresponding entries together. So the matrices must have the same size m by n, like the vector addition. The next important operation is the Matrix Multiplication.


      앞의 행렬의 열이 개수와 뒤의 행렬의 행의 개수가 같아야 한다는 조건이 필요합니다. 즉, 2 by 3 행렬에 3 by 2 행렬을 곱할 수 있습니다. 앞의 행렬의 열(column)의 개수, 하나, 둘, 셋과 뒤의 행렬의 행(row)의 개수, 3이 일치할 때 곱할 수 있습니다. 왜냐하면 행렬곱은 첫 번째 행렬의 행(row)과 두 번째 행렬의 열 (column)을 곱해주기 때문에 그렇습니다. a하고 g, 두 번째 b하고 i, 세 번째 c 하고 k가 곱해서 더한 것이 첫 번째 (1, 1) 성분이고, 첫 번째 행(row)과 두 번째 열(column)을 곱하게 여기이고, 그리고 두 번째 행(row)와 첫 번째 열(column)을 곱한 게 (2, 1) 성분(component)이고, 두 번째 행(row)와 두 번째 열(column)을 곱한 게 (2, 2) 성분(component)가 되는 식의 계산입니다. 기역자로 곱해지죠. 기역! 기역, 그래서 이를 <행렬곱셈에서의 King 세종의 법칙>이라고도 얘기합니다.


    For matrix multiplication, the number of columns in the first matrix must be equal to the number of rows in the second matrix. For example, we can multiply a 2 by 3 matrix by a 3 by 2 matrix. This is because i-th row of the first matrix A and j-th column of the second matrix B make (i, j)-th entry of AB. For example the first (1, 1) entry of AB is ag+bi+ck, and (2, 2) entry of AB is dh+ej+fl.

    Since this multiplication is done with the shape of the character <ㄱ> which is the first Korean alphabet ㄱ! So we call it <the King Sejong’s Rule for the Matrix Multiplication> in Korea.


     이 행렬과 이 행렬이 있으면 첫 번째 (1, 1) 성분은 첫 번째 행(row)와 두 번째 열(column)이 곱해져서 22가 나오고, 두 번째, (2, 2) 성분은 두 번째 행(row)와 두 번째 열(column)을 곱해서 얻어집니다.


    With these two matrices, the first (1, 1) entry of AB becomes 22 after multiplying the first row of A by the first column of B, and (2, 2) component is the product of the second row of A and the second column of B.



    다음 행렬의 연산법칙. 행렬 가 적당한 사이즈의 행렬이고, 가 스칼라일 때, 다음과 같이 덧셈의 교환법칙이 성립되고, A+B=B+A이고, 덧셈에 대한 행렬의 결합법칙이 성립되고, A에 B, C 곱한 행렬을 곱한 것과 AB를 곱한 행렬에 C를 곱한 것이 같다는 곱셈의 결합법칙이 행렬 사이에 성립하고, 또 분배법칙이 양쪽으로 성립하고, 스칼라의 분배법칙이 성립합니다. 여러분들이 대개 보아왔던 내용들이죠. 실수 벡터들의 성질들하고도 비슷하죠. 일반화된 모양들입니다. 행렬에다가 상수 1을 곱하면 언제나 자기 자신입니다.

    [Law of Matrix Operations] Suppose that the matrices are appropriately sized matrices and are scalars,

    (1) A+B=B+A, the commutative law of matrix addition

    (2) means the associative law of matrix addition works.

    (3) means the associative law of matrix multiplication works.


    (5)-(9) means usual distributive law holds too.

    (10) If you multiply a matrix by the constant 1, it is itself.


    The above rules are similar to the rules what we've have seen for vectors. We can think this rules are a generalized form of the Rules for vectors because a column vector is a n by 1 matrix.


    다음 행렬에 대하여 , , 를 계산하라는 문제를 보시면, A, B, C가 주어졌을 때, A+B는 성분들을 더해주면 되고, 상수배 해주는 것은 성분별로 2배해주면 되고 AC를 구할 때 첫 번째 행(row)와 첫 번째 열(column)을 곱한 이것이 (1, 1) 성분이 됩니다.


    For given matrices A, B, and C, A+B is obtained from adding the corresponding entries of A and B, the scalar multiplication 2A is obtained from just doubling for each entry of A, and for evaluating AC, we use the King Sejong's Rule.


    1 곱하기 3 + 2 곱하기 0 + –4 곱하기 7 하면 –25가 되겠죠.


    (AC)_{1,1} will be  1x3 + 2x0 + (-4)x7 = -25.


    그리고 (2, 2) 성분은 10이 나오는데, 두 번째 행(row)와 두 번째 열(column)을 각각 성분별로 곱해서 더해준 결과입니다. 그것을 코드를 이용해서 해보시면, A행렬을, B행렬을, C행렬을 다음과 같이 정의하고, A, B를 더해서 프린트하라, 그 다음에 두 배 A를 프린트하라, A와 C 곱한 것을 프린트하라고 하면, 위에서 손으로 계산한 것과 같이 A, B 행렬, 2A 행렬, 그리고 A, C를 곱한 행렬이 정확하게 일치하게 나오는 것을 확인할 수가 있습니다.


    And the (2, 2) entry (AC)_{2,2} is 10, which is the result of (entry wise) multiplying the second row of A and the second column of C and adding them all. We try to do it with Code.
       A = matrix([[1, 2, -4], [-2, 1, 3]])

    B = matrix([[0, 1, 4], [-1, 3, 1]])

    C = matrix([[3, -1], [0, 5], [7, 1]])

    print(A + B)

    print(2*A)

    print(A*C)


    Then, we can see that the matrix A, B, 2A and AC come out exactly same as those computed by hand.


    행렬과 스칼라를 임의로 생성하여 행렬의 덧셈과 실수배, 곱셈을 다음과 같이 확인할 수가 있습니다. 이 부분은 2차시에서 이어서 학습하도록 하겠습니다.


    By randomly generating matrices and scalars, we can check our matrix computations as follows. We will continue to study <Matrix operations> in the next lecture.


    지금까지 우리가 데이터로부터 행렬을 정의하고, 그리고 행렬의 연산을 학습하였습니다.                                              [2주 1차시 끝]


    So far, we have defined a matrix from a data and learned related matrix operations. [End of 1st lecture (Week 2)]



    [2-2 pre]

    2주의 2차시에서는 Chapter 2 '데이터와 행렬' 중 <2.4절> 행렬 연산과 <2.5절> 행렬의 연산 법칙에 대하여 자세하게 살펴보도록 하겠습니다.


    In this second lecture, we will take a closer look on Section <2.4> Matrix operations and Section <2.5> Properties of  Matrix operations.



    [2주차 2강]  반갑습니다. K-MOOC. 인공지능을 위한 기초수학 입문. 2주차 2차시. '데이터와 행렬'입니다. 지난 2주차 1차시에서는 데이터로부터 행렬을 생성하고 그 행렬들 사이의 연산을 정의했습니다.


    [Week 2, Lesson 2] Welcome to <K-MOOC. Introductory Mathematics for Artificial Intelligence> class. Now we will start the second session of Lecture 2, 'Data and Matrix'. In the first session, we generated a matrix from the given data and defined matrix operations.


    오늘은 이어서 실제 행렬계산들을 해보도록 하겠습니다.

    Today, we do matrix computations with codes.


    지난 시간에 두 개의 행렬이 주어지면 행렬사이의 덧셈합(vector addition)과 스칼라배(scalar multiplication) 및 행렬곱셈(matrix product) 연산을 배웠습니다. 특히 행렬곱셈(matrix product)에서 다음과 같이 <King 세종의 법칙 (기역자 법칙)>, 즉 앞의 행렬의 i 번째 행(row)과 두 번째 행렬의 j 번째 열(column)을 순서대로 쌍으로 곱해서 모두를 더한 것이 곱한 행렬 AB 의 (i , j) 성분이 된다는 것을 학습했습니다.


    In the last lecture, we learned the vector addition, scalar multiplication, matrix multiplication, and their properties. In the matrix multiplication called <the King Sejong's rule (Letter 'ㄱ's rule)>, we learned that the (i, j)-th entry of the matrix product AB is the inner product of the i-th row of A and the j-th column of B.


    그 내용을 실제 코드로 실습해보도록 하겠습니다.

    Let's do matrix computations with python-based codes.


    이번 코드 A = random_matrix(ZZ, 4, 5, x = -10, y = 10) 는  -10부터 9까지의 정수를 성분으로 하는 4x5 행렬 A를 임의로 생성하고, 4x5 정수 행렬의 성분(entry)을 정수 중에서도 –10과 9 사이에서 random 하게 생성한다는 의미입니다.


    The code 'A = random_matrix(ZZ, 4, 5, x = -10, y = 10)' will generates a 4x5 matrix A whose entries are integers randomly chosen from -10 to 9.


     두 번째 B = random_matrix(ZZ, 4, 5, x = -5, y = 5) 로는 B 행렬을 어떻게 생성하냐면, 정수에서 4x5 행렬을 –5에서 4 사이에서 random하게 생성하고, 그 다음 C 행렬 C = random_matrix(ZZ, 5, 3, x = -5, y = 5)은 정수에서 5x3 행렬을 생성하는데 성분은 -5와 4 사이에서 생성 하자는 의미입니다. 스칼라 k = ZZ.random_element(x = -5, y = 5) 도 생성합니다. 여기서 5라는 의미는, 실제로 5라는 의미는 0을 하나의 숫자로 표현하기 때문에 실제 표현되는 숫자는 이 숫자 5보다 하나 작은 숫자 4 까지가 실제 행렬에 스칼라 성분(entry)를 생성할 때 사용합니다.


    And the code 'B = random_matrix (ZZ, 4, 5, x = -5, y = 5)' will generate a 4x5 matrix whose entries are integers randomly chosen from –5 to 4.

    The code  'C = random_matrix (ZZ, 5, 3, x = -5, y = 5)' will generate a 5x3 matrix with the integer randomly chosen from -5 to 4. The code ' k = ZZ.random_element(x = -5, y = 5)' will generate a scalar k randomly chosen from -5 to 4. Since the index starts with 0 in python/sage, 5 in the code above means the number 4.


    랜덤(random) 하게 생성한 4x5 행렬을 프린트하고, 4x5 행렬 B를 프린트하고, 5x3 행렬 C를 프린트하고, 스칼라 k를 프린트 한 후에, 그것을 둘을 합해봐라, 덧셈, 행렬의 덧셈. 스칼라 배. 그 다음 행렬곱(product), A 행렬과 C 행렬의 행렬곱(product)를 계산하시오 하는 명령을 주었습니다.


    Print a randomly generated matrix A, B, C and scalar k, then add A and B and kA. Then we do the matrix multiplication of matrix A and C by codes.


    그럼 random하게 4x5 행렬 A, B가 생성됐고, 그 다음에 C 행렬이 생성됐고. k가 –1로 생성이 됐고, 그 명령어를 한 번에 클릭하면, A+B, k배의 A, A와 C를 곱한 행렬의 product가 바로 계산이 됩니다.


    Then, two 4x5 matrices A and B were randomly generated, followed by C and k as –1. If we click the ‘evaluate’ button, then A + B, k times A, and AC will be immediately evaluated.


    여러분들이 지금 확인해본 이 연산을 통하면, 단지 4x5 행렬이 아니라 10x20 행렬 또는 50x100 행렬, 555x999 행렬 같은 것 사이에 연산도 얼마든지 가능하다는 그런 의미입니다.

     

    Using similar codes, we can handle not only 4x5 matrices, but also 10x20 matrices, 50x100 matrices, 555x999 matrices, and so on.


    열린문제 한번 해볼까요? 위에서 배운 벡터와 행렬에 대한 연산을, 여러분들은 어떤 벡터나 어떤 행렬을 주더라도 연산을 하실 수 있게 됐으니까. 다른 교재에 나온 복잡한 그런 행렬들의 합이나 곱이나, 스칼라 배 등을 실제 실습실에서 해보시고 그 내용을 QnA 문의게시판에서 공유해 보시기를 바랍니다.


    Let's consider an open problem. Since we can perform operations on any vectors or matrices learned above, you may try to do vector addition, scalar multiplication, matrix multiplication of such large-sized matrices from other textbooks with our cyber-lab and share what you’ve done in the QnA board of LMS.


    예제는 다음과 같습니다. 수면시간, 운동시간, 칼로리 섭취량이 체중과 혈압에 미치는 영향을 행렬로 표현해보겠습니다.


    Here is an example: let's use a matrix to explain a relationship between sleeping hours, exercise hours, calorie intake and the weight and blood pressure of a person.


     벡터는 행렬, 행벡터라고 부르고, 또는 행렬을 열벡터라고 생각할 수 있습니다.


    A n-dimensional vector is called a matrix or a row vector, and an matrix can be thought of as a column vector.


    즉, 벡터는 행렬의 스페셜 케이스라고 이해할 수 있습니다. 수면시간, 운동시간, 칼로리 섭취량에 따라서 그것이 사람의 체중과 혈압에 영향을 미치니까. 이 그림의 아이디어는 고려대 영문과 남호성 선생님이 소개하신 것을 활용했습니다. 영문과에서 인공지능을 가르치고 배운다니 신기하죠. 인공지능에 필요한 수학이라면 당연히 수학 교수님들이 설명해줘야 할 것인데, 영문과 교수님이 인공지능에 필요한 수학을 강의 하신다는데 자극을 받아서, 그 때부터 공부해서 제 전공인 행렬과 관련된 <인공지능에 필요한 수학>에 대한 설명을 제가 하게 되었습니다.


    In other words, a vector can be understood as a special case of a matrix. Because sleeping hours, exercise hours, and calorie intake may affect a person's weight and blood pressure. This .\PICture was introduced by Prof. Ho-Sung Nam at Korea University.


    수면 시간이 , 운동시간이 , 칼로리 섭취량이 인 사람인 경우 ‘체중은  , 혈압은 가 대략 되어야 정상이다.’라는 함수를 생각할 수 있습니다. 그러면 벡터를 입력(input)하면 (행렬을 곱하여) 벡터가 나오게(output) 되니까 중간에 행렬이  involve 됩니다. 그래서 우리반 수동이의 하루 평균 수면시간과 운동시간 그리고 칼로리 섭취량에 따라 체중과 혈압이 변하는 상황을 조사하여 만든 초기 행렬 를 다음과 같이 구했다고 합시다. 만일 수동이가 요즘 하루 평균 수면시간 8 시간, 운동시간 1 시간, 칼로리 섭취량이 1500 kcal을 유지해 왔다면, 주어진 행렬 를 이용해 계산한 예측값이 무엇이 될까하는 것은 간단합니다. 이 행렬에다가 수동이의 활동 정보를 곱해주면 체중은 71.9 kg 또 혈압은 115.83 mmHg가 되는 것을 행렬 A를 정의하고, 벡터 x를 정의해서 x에다 A를 곱해주면 다음과 같이 체중과 혈압이 나옵니다.

     

    In case of a person with sleeping hour , exercise hour , and calorie intake , ‘it would be normal if his weight is and his blood pressure is in some degree.’ If one was an input (vector) and the other was an output (vector), then we can think of a function in between. When a function takes a vector as an input, and give an output as a vector, then this function can be considered as a matrix. Suppose we have a matrix by mathematical modeling for the above situation. We have a health data on a student named 'Sudong' in our class. If Sudong has maintained an average sleep time of 8 hours per day, exercise time of 1 hour, and calorie intake of 1500 kcal a day for a certain period, it is simple to determine what the predicted value will be given by multiplying the given matrix . If this matrix is ​​multiplied by the health information of Sudong, the expected weight is 71.9 kg and the expected blood pressure is 115.83 mmHg. If we define the matrix A and the vector x, and ask to find Ax, then we will have the expected weight and blood pressure of him as follows.


    여기서는 Ax=y 대신 x^T A^T = y^T 와 같이  x^T 를 먼저하고 A^T 를 나중에 곱해줬습니다.


    Instead of Ax=y, we used the equation x^T A^T = y^T. We may note that this is a usual notation that data scientists are using now in AI.


    수학에서는 열(column) 벡터를 집어넣고 열벡터를 나오게 하는데, 통계라든지 인공지능 또는 공학에서는 행(row) 벡터를 집어넣어 행(row) 벡터를 나오게 하기 때문에, 연산이 Ax=y 대신에 전치행렬 개념을 사용해서 행(row) 벡터를 집어넣고 행렬을 곱하면 행(row) 벡터가 나오는 식으로 바꿔줘서 코드가 이렇게 바꿔진 겁니다. 표현만 다르지 원리는 차이가 없습니다.


    In mathematics, a column vector is used for input and output. But in statistics, artificial intelligence, or engineering, a row vector is used for input and output, so the code is changed in this way by taking the transpose of the equation x^T A^T = y^T. A row vector is multiplied by a transpose matrix, and gets a resulting row vector. It is only different in the representation, not in the principle.


    그래서 수동이의 예상되는 체중과 혈압은 체중 71.9 kg, 혈압은 115.83 mmHg(혈압이나 안압을 재는 수은주 밀리미터) 과 같다는 설명을 얻게 된 것입니다.


    So, it was explained that the expected weight and blood pressure of Sudong was 71.9 kg and 115.83 mmHg.


    이 부분에 대해서 제가 더 설명해 놓은 게 있는데, 인공지능, 수학으로 타파하자. 딥러닝에 대한 설명을 다음 주소에서 소개했는데, 여기를 참고해서 보시면 필요한, 중학생들을 대상으로 한 내용인데, 참고로. 지금보시면 ‘주니어폴리매스’ 인공지능, 수학으로 타파하자에서 딥러닝을 간단히 소개한 내용입니다.


    I have written a series of articles “Go Through AI with Mathematics” for Junior High students in Math DongA journal. The explanation on deep learning is introduced at the following link. Please refer to this for a brief introduction on deep learning in 'Junior Polymath', which was written for middle school students.


     인공지능의 아버지 알렌 튜링(Turing)에 대해서 소개도 했고, 벡터덧셈과 스칼라 곱(multiplication)을 소개했습니다.


    I introduced Alan Turing, the father of artificial intelligence, and vector addition and scalar multiplication.


    행렬이 실제 덧셈, 뺄셈으로 계산되는 여러분들이 행렬이 바꿔지면 바꿔지는 행렬에 대해서도 자연스럽게 연산이 바꿔진 걸 확인해볼 수 있습니다. 스칼라 배도 마찬가지입니다. 중학생들도 따라서 하게 만들었습니다. 체중과 혈압에 대한 이 예를 여러분들 교안으로 사용했으며, 그렇게 하기 위해서 벡터는 어떻게 정의하는지, 행렬은 어떻게 정의하는지, 행렬곱은 어떻게 정의하는지 식으로 소개를 하였습니다.


    We can practice some matrix operations at here. It can be easily done by junior high school students as well. I used this example of weight and blood pressure in our textbook, in order to introduce how to define a vector, a matrix, and matrix multiplication.


    만일 실제 측정한 값이 체중 72 kg, 혈압 116 mmHg이였는데, 예측값과 실제측정값 사이에 차이가 나는 이유는 데이터가 모자라서 행렬 가 부정확한 것이 이유였다면, 이번에 측정한 데이터를 활용하여, 예측값이 측정한 값과 같아지도록 행렬 의 성분들을 약간씩 조정하여 가 되는 수정된 행렬 를 찾을 수가 있습니다. 이것이 딥러닝의 <오차 역전파법> 아이디어입니다.


    If the observed value of body weight and blood pressure has a significant difference with the predicted value even though Sudong was normal, the matrix must be modified. We can find the modified matrix such that by slightly adjusting the components of the matrix to match the predicted value with the observed values. This is the idea of <Back propagation> in deep learning.


    즉, 우리가 행렬을 만들어놨고, 많은 그룹의 학생들 데이터를 입력해서 예측한 것이 대충 맞아떨어져야 되는데, 만일 우리가 기대한 만큼 맞아 떨어지지 않는다고 하면, 그 함수, 즉 행렬을 계속 데이터를 활용해서 수정해가서 대략 예측하는 값이 나오는 함수로 바꿔주는 그 중간 과정이 오차 역전파법의 아이디어입니다. 이후 13주차와 14주차에 part 4에서 자세히 설명해드리도록 하겠습니다.


    Suppose that we have an initial matrix. If the predicted values obtained from this matrix and the given data is different from the observed values, then we should modify to find a better matrix. This is the idea of Back propagation in artificial neural networks. We will explain in detail later in Lecture 13.


    다음은 전치행렬, 아까 전치행렬을 간단히 보신 셈인데. 행렬 에 대해서 행과 열을 바꾸어 얻어진 행렬을 A의 전치행렬(transpose)라고 하고, 예를 들어서 가 행렬이면 A의 전치행렬은 행렬이 되는 겁니다. 전치행렬 A^T 라는 것은 는 로 행렬의 행(row)와 열(column)을 바꾸는 것입니다.


    The transpose of a matrix is simply a flipped version of the original matrix. We can transpose a matrix by switching its rows with its columns. For example, if A is a matrix, then A^T is a matrix.


    전치행렬 A^T을 구할 때, A의 은 으로, 는 으로, 은 으로, 은 로, 는 로 이렇게 바꿔주면 되겠죠. 이렇게 은 로. 은  으로, 는 으로, 은 으로 이런 식으로 성분이 행과 열이 바꿔서 얻어진 행렬이 전치행렬 A^T입니다.


    And (i, j) entry of A^T is the (j, i) entry of A.


     행렬 와 임의의 스칼라 에 대하여 전치행렬은 다음 식을 만족합니다. 중요한 성질입니다.


    The following properties of a transpose matrix hold.


    행렬 A의 전치행렬(transpose)의 전치행렬(transpose)은 자기자신이다, A+B transpose는 A의 transpose 더하기 B transpose다. 중요한 것은 이겁니다. AB의 transpose는 B transpose 곱하기 A transpose 이다.  이것은 쉽게 증명할 수 있습니다. 궁금하면 물어보세요. 답 설명해드릴게요.


    (1) (A^T)^T = A.  The transpose of the transpose of A is A

    (2) (A+B)^T = A^T + B^T. The transpose of a sum is the sum of transposes.

    (3) (AB)^T = B^TA^T.  The transpose of a product is the product of the transposes in the reverse order.


    상수배한 것의 transpose는 transpose에 상수배를 해준 거 즉, 상수를 곱해서 transpose를 구한 것은 transpose를 먼저 한 다음에 상수배 한 것과 같다는 것입니다.

    (4) (kA)^T = kA^T.  The transpose of a scalar multiple of a matrix is the  scalar multiple of its transpose matrix.


    다음 행렬들의 전치행렬을 각각 구하시오. A와 B가 주어지면 A의 transpose와 B의 transpose를 손으로 이렇게 구할 수 있지만, 행렬 A를 정의해주고, B를 정의해준 다음에, A transpose, B transpose를 바로 우리가 print out 해서 이렇게 볼 수가 있습니다. 위의 코드를 바로 활용한 것입니다.


    [Example] Find the transpose matrix for each of the following matrices. For given A and B, we can obtain the transpose of A and the transpose of B by using the following codes.


    그 다음은 대각선 행렬이죠. 대각선성분 외에는 모두가 다 인 특수한 모양의 행렬입니다. 앞으로도 많이 사용됩니다. 그 중에 특수한 예가 단위행렬(identity matrix) 또는 항등행렬이라 하는데요. 주대각선성분만 1이고 나머지는 다 0인 행렬입니다. 대각선 행렬은 간단하게 다른 성분이 다 0이기 때문에 대각선 성분만 쭉 쓰고, 앞에 diag라고 써줘서 표현하기도 합니다. 단위행렬(identity matrix, 항등행렬)은 차의 단위행렬 로 이렇게 표시합니다. 주대각선성분이 모두 1인 대각선 행렬을  단위행렬(identity matrix)이고, 다음이 성립합니다. 즉 행렬의 곱셈에 대한 항등원 역할을 한다. 즉, 행렬 A에다 단위행렬을 곱하면, 왼쪽에 곱하건 오른쪽에 곱하건 다, A 자기 자신이 된다.


    A diagonal matrix is a matrix in which the entries outside the main diagonal are all zero. As an example, the identity matrix is a square matrix that has 1's along the main diagonal and 0's for all other entries. Multiplying any matrix by the identity results in the matrix itself.


    삼각행렬은 하삼각행렬과 상삼각행렬을 얘기할 수 있는데, 상삼각행렬은 위쪽에만 성분이 있고 나머지는 모두 0인, 하삼각행렬은 아래쪽에 성분들이 주로 있고 주대각선 성분 위는 모두 0인, 모양을 보고 아래쪽이 0이면 상삼각행렬, 위쪽이 0이면 주대각선성분 위가 다 0이면 하삼각행렬이라고 부릅니다.


    A triangular matrix is a special kind of square matrix. A square matrix is called lower triangular if all the entries above the main diagonal are zero. Similarly, a square matrix is called upper triangular if all the entries below the main diagonal are zero.


    그리고 대칭행렬은 A의 전치행렬 (A transpose)가 A하고 같을 때, 대칭행렬이라고 합니다. 이런 행렬이 대칭행렬의 예입니다. A하고 A의 전치행렬(transpose)이 같은 이런 행렬은 굉장히 중요합니다.


    A symmetric matrix is a square matrix that is equal to its transpose. Symmetric matrices will play an important role in this class.


    역행렬, 정사각행렬이 아래 성질을 만족하면, 아래성질을 만족하는 행렬 가 존재하면 가 가역행렬(invertible 행렬)이라고 얘기합니다. 즉, 행렬 A에 대해서 A에 다른 행렬 B를 곱했더니 항등행렬(identity matrix)이 되고 또 A에다 B를 왼쪽에 곱해도 항등행렬(identity matrix)이 되는 행렬 B가 존재하면 우리는 가역행렬(invertible matrix)이라고 하고, 를 의 역행렬(inverse matrix)이라고 하면서 ‘A inverse’ 로 표현합니다. 어떤 행렬의 역행렬은 유일하게 존재하므로 이런 행렬 B를 A inverse ( A^{-1} )로 쓸 수 있습니다. 이런 역행렬 B가 존재하지 않을 때 우리는 비가역행렬(noninvertible) 또는 특이행렬(singular matrix)이라고 합니다.


    An n-by-n matrix A is called invertible (also nonsingular) if there exists an n-by-n square matrix B such that AB = I = BA.  If this is the case, then the matrix B is uniquely determined by A and is called the inverse of A, denoted by A^−1.

    A square matrix that is not invertible is called singular (noninvertible), that is. there is no B, such that AB = I = BA.


     차의 정사각행렬 가 가역이고 가 이 아닌 스칼라일 때, 다음이 성립 합니다. <A가 invertible이면 A의 역행렬(inverse)도 가역이고 A inverse의 inverse는 자기 자신이다>.


    For invertible matrices, the followings hold.

     If A is invertible, then A-1 is also invertible and (A^-1)^-1 = A.


    AB가 가역행렬(invertible matrix)이면, AB의 역행렬(inverse)는 B inverse 곱하기 A inverse입니다. 굉장히 중요한 성질입니다.


    If A and B are invertible matrices, then AB is also invertible and (AB)^-1 = B^-1A^-1. This property is very important.


    상수배에서도 다음과 같은 성질이 성립하고 전치행렬(transpose)도 중요한 성질인데, ‘A의 전치행렬(transpose)의 역행렬(inverse)은 A의 역행렬(inverse)의 전치행렬(transpose)과 같다’입니다.


    If A is invertible and k is a nonzero scalar, then kA is also invertible and (kA)^-1 = 1/k A^-1. And if A is invertible, then A^T is also invertible and (A^T)^-1 = (A^-1)^T.


    다음에는 ‘역행렬을 구하시오’하는 문제가 나오는데, 2차 행렬의 역행렬을 구하는 것은 간단합니다. 그래서 A 행렬이 이렇게 주어지면 A inverse는 다음과 같이 구해질 수 있습니다. 행렬 A를 정의하고, 명령어로 A의 역행렬(inverse)를 구하면 됩니다. 명령어 A.inverse() 로 행렬 A의 역행렬을 구하라고 하면 바로 구해줍니다.


    [Example] Find the inverse matrix.


    It is simple to find the inverse matrix of a 2x2 matrix. So for a given matrix A, the inverse of A can be obtained as follows: define matrix A, then use the Code to find the inverse of A. If we use the code 'A.inverse()' to find the inverse of matrix A, we will have the inverse.


    중요한 것은 이 명령어가 2차 행렬뿐만 아니라 3차 행렬, 4차 행렬, 5차 행렬, 10차 행렬까지 작동되는 것입니다. 그래서 3차 행렬을 이렇게 주고 행렬 A가 invertible하냐고 A.is_invertible() 로 물어보면, ‘이 행렬은 invertible하지 않습니다’라고 나옵니다. 역행렬이 존재하지 않는다는 것이죠.


    The important thing is that this command works not only for 2x2 matrices, but also works even for large-sized matrices, for example, 10x10 matrices. So, if we give a 3x3 matrix like this and use 'A.is_invertible()' to determine whether the matrix A is invertible or not, it may say 'this matrix is not invertible'. It means, there is no inverse matrix.


    그래서, 열린문제에서는, ‘인터넷이나 다른 교재에서 5차 (이상) 행렬을 찾아서, 전치행렬과 역행렬이 존재하는지를 확인하고, 존재하면 찾아보십시오’하는 것을 2주차 과제로 하겠습니다.


    Our open problem of this lecture is to ‘find the 5x5 (or bigger sized) matrix in other textbook, and check whether the transpose and inverse matrices exist, and find them if it exists.’


    그래서 행렬연산에 대해서 본인이 요약하고, 실습하고, 질문하거나 답변한 내용, 그 과정에서 느낀점, 깨우친 점을 코멘트로 달아보도록 하십시오. 역행렬과 관련해서는 참고자료로 K-MOOC 선형대수학 주소에 들어가시면 참고해서 보실 수가 있으실 것입니다.


    And we may summarize what we did in practice, or learned from today’s lecture. Please visit the following link for reference:

      http://matrix.skku.ac.kr/LA/

      http://matrix.skku.ac.kr/K-MOOC-LA/index.html .


    [2-2 Review]


    여러분 수고하셨습니다. 우리 2주차 데이터와 행렬에서는, 벡터, 또 데이터들을 벡터로, 여러 데이터들을 행렬로, 그런 행렬들에 색깔을 주어서 이미지로, 또 다른 사회의 데이터들을 행렬과 텐서로 표현하는 process를 설명했고, 그 내용을 1절에서는 순서쌍, 순서조와 벡터, 벡터들 사이의 연산, 그 다음에 행렬과 텐터를 소개한 후에, 행렬 연산과 행렬의 연산법칙에 대해서 학습을 하였습니다. 그 내용을 week2 주소에서 확인할 수 있었습니다.


    In this week, we learned about data, vector, matrix, and tensor. We learned that a photo image can be written as a matrix. In Section 1, we learned ordered pairs, n-tuples, vectors, vector operations, matrix operations and the rules on matrix operations. All can be found in http://matrix.skku.ac.kr/math4ai-intro/W2/.


     배운 내용들을 복습하면, 주어진 데이터로부터 데이터를 벡터로 표현하고, 벡터들의 시각적 이미지를 소개한 후에, 우리 사회에 사용되고 있는 이미지와 사진과 어떻게 관계가 있는지, 행렬과 어떻게 관계가 있는지를 설명한 후에, 행렬 사이의 연산 특히 행렬 곱과 전치행렬 및 역행렬에 대해서 학습했습니다.


    Let’s review today’s lecture. A data can be represented as a vector or a matrix. Many real data including images can be written as a matrix. We learned matrix operations, including the inverse and the transpose of a matrix.


    그리고 행렬곱은 다음과 같이 두 행렬이 있으면 앞의 행렬의 행(row)하고 두 번째 행렬의 열(column)을 곱한 것이 곱해진 행렬의 성분이 되도록 그래서 행렬 C의  성분은 A의 번째 행(row)와 B의 번째 열(column)을 성분끼리 곱해서 더해서 표현되는 식으로 설명해 드렸습니다. 그리고 주대각선 성분들을 더한 것이 trace라는 것, 역행렬 및 역행렬을 구하는 방법과 역행렬의 성질을 소개했고, 특수행렬로 상삼각행렬과 하삼각행렬을 소개하였습니다.


    We have learned the matrix multiplication and the King Sejong's 'ㄱ' rule to finding (i,j)-th entry of AB. And the properties of an inverse matrix were introduced. And the upper and lower triangular matrices were introduced as special matrices.


    이어서 3주차에는 데이터의 분류(classification)에 대해서 학습하도록 하겠습니다.


    Next, in Week 3, we will learn about <Classification of data>.


    감사합니다. Thank you.




           그림입니다.
원본 그림의 이름: 200605170651_poster.jpg
원본 그림의 크기: 가로 3275pixel, 세로 4631pixel

     

    Week 3. Introductory Mathematics for Artificial Intelligence

     

                인공지능을 위한 기초수학 입문

     

    3.  데이터의 분류 38

    3.1 데이터의 유사도

    3.2 거리

    3.3 노름(norm, 크기)

    3.4 노름(크기, 거리)을 활용한 데이터의 유사도 비교

    3.5 사잇각을 활용한 데이터의 비교

    3.6 코사인 유사도의 개념

    3.7 내적

    3.8 사잇각

    3.9 코사인 유사도의 계산


    [3-1 pre] 여러분 반갑습니다. 이제 3주차 <데이터의 분류>에 대하여 학습하도록 하겠습니다. 1차시에는 먼저 데이터를 분류하는 ‘유사도’에 대하여 학습합니다.


    [3-1 pre]  We will study <classification of data> in Week 3. First of all, we now introduce the concept of <Similarity in data>.


    데이터의 특징에 따라 데이터가 어느 범주에 속하는지 판단하는 것을 우리는 분류(classification)하고 합니다. 이를 위해서는 데이터 들이 서로 얼마나 가까운지 즉 유사한지를 판단할 수 있어야 하는데, 먼저 거리를 이용하여 데이터가 얼마나 유사한지 판단하는 거리 척도와 그를 위한 벡터의 노름(norm)에 대하여 학습합니다. 그리고 데이터의 패턴, 방향에만 관심이 있는 경우 사용할 수 있는 척도인 코사인 유사도와 내적(inner product)에 관하여 학습합니다.


    Classification is the problem of identifying to which categories a new data belongs, based on the characteristics of the given data. To do classification, we need to determine how close each other between two different data, that is, we need a distance measure. Now we introduce a concept of distance on vectors. The 'distance similarity' can be used to tell how similar/close the given data are. It can be done with 'the definition of norm'. Also we will learn about 'cosine similarity' and 'inner product.’ 'Cosine similarity' can be used when we only interest in patterns and directions of data (or vectors).


    설명의 순서는 데이터의 유사도, 거리(distance), 거리의 일반화된 개념인 노름, 그리고 노름을 이용한 데이터의 유사도 비교, 사잇각을 활용한 데이터의 비교, 코사인 유사도, 내적과 사잇각을 소개한 후에, 코사인 유사도를 직접 계산하여, 같은 성질을 가진 데이터들을 분류(classify)하는 데이터의 분류를 학습하도록 하겠습니다.


    We will study them in the following order.

    1) Similarity of a data

    2) Distance

    3) Norm, that is, a generalization of distance

    4) Similarity between data using norm

    5) Cosine similarity

    6) Inner product and Classification


    사람 A와 사람 B의 유사도를 어떻게 평가할 수 있을까? 누가 누구하고 더 친한가를 거리를 이용하여 확인할 수 있고, 방향을 이용해서도 분류할 수 있습니다.


    How can you evaluate the similarity between the person A and the person B? We can check who is more friendly with whom by distance, and we can also do a classification by the other measure such as direction/cosine similarity.


    이것을 학습하기 위해서 우리는 내적이라는 개념도 소개합니다.


    In order to do so, we need 'the concept of inner product’

    [3-1 pre] 여러분 반갑습니다. 이제 3주차 <데이터의 분류>에 대하여 학습하도록 하겠습니다. 1차시에는 먼저 데이터를 분류하는 ‘유사도’에 대하여 학습합니다.


    [3-1 pre]  We will study <classification of data> in Week 3. First of all, We start to introduce the concept of <Similarity in data>.



    데이터의 특징에 따라 데이터가 어느 범주에 속하는지 판단하는 것을 우리는 분류(classification)하고 합니다. 이를 위해서는 데이터 들이 서로 얼마나 가까운지 즉 유사한지를 판단할 수 있어야 하는데, 먼저 거리를 이용하여 데이터가 얼마나 유사한지 판단하는 거리 척도와 그를 위한 벡터의 노름(norm)에 대하여 학습합니다. 그리고 데이터의 패턴, 방향에만 관심이 있는 경우 사용할 수 있는 척도인 코사인 유사도와 내적(inner product)에 관하여 학습합니다.


    Classification is the problem of identifying to which of a set of categories a new data belongs, based on the characteristics of the given data. To do classification, we need to determine how close each other between two different data, that is, we need a distance measure. Now we introduce a concept of distance on vectors. The 'distance similarity' can be used to tell how the given data are similar. It can be done with 'the definition of norm'. Also we will learn about 'cosine similarity' and 'inner product.’ 'Cosine similarity' can be used when we are only interested in patterns and directions of data (or vectors).



    설명의 순서는 데이터의 유사도, 거리(distance), 거리의 일반화된 개념인 노름, 그리고 노름을 이용한 데이터의 유사도 비교, 사잇각을 활용한 데이터의 비교, 코사인 유사도, 내적과 사잇각을 소개한 후에, 코사인 유사도를 직접 계산하여, 같은 성질을 가진 데이터들을 분류(classify)하는 데이터의 분류를 학습하도록 하겠습니다.


    We will study them in the following order.

    1) Similarity of a data

    2) Distance

    3) Norm, that is, a generalization of distance

    4) Similarity between data using norm

    5) Cosine similarity

    6) Inner product and Classification


    사람 A와 사람 B의 유사도를 어떻게 평가할 수 있을까? 누가 누구하고 더 친한가를 거리를 이용하여 확인할 수 있고, 방향을 이용해서도 분류할 수 있습니다.


    How can you evaluate the similarity between person A and person B? You can check who is more friendly with whom by distance, and you can also do classification by the direction similarity.



    이것을 학습하기 위해서 우리는 내적이라는 개념도 소개합니다.


    In order to do so, we need 'the concept of inner product’


    3.1절 데이터의 유사도.


    Sec 3.1 Similarity of data


    인공지능이 데이터를 바탕으로 수행하는 주요 업무는 크게 판단(decision)과 예측(prediction) 으로 나눌 수 있습니다. 예를 들어, 이미지 인식으로 읽은 사진으로부터 사진 속의 사람이 누구인지 판단하고, 음성인식으로 저장된 음성신호로부터 말하는 내용과 의미를 판단하고, 의료영상을 인식한 후 스스로 비교 분석하여 질병의 유무와 정도를 판단합니다. 그리고 전자상거래의 인터넷 기록을 활용한 사용자의 과거 구매 이력으로부터 그 사용자가 어떤 상품에 관심을 가질지 예측하여 다음 구매할 상품을 추천하는 일, 금융에서 주식의 과거 가격과 거래 정보로부터 미래 가격의 흐름을 예측하는 일, 바둑 대국의 현재 상황에서 어떤 위치에 다음 바둑돌을 놓을 때 승률이 어떻게 변할지 예측하는 것. 이런 일들이 모두 인공지능이 잘 하는 일들입니다.

     

    The two important tasks on data that artificial intelligence performs are  'decision and prediction'. For example, AI determines who the person is in a given .\PICture by image recognition, determines what is said and its meaning in a given voice signal by voice recognition, and compares and analyzes the medical image to determine the presence and extent of the disease. And it predicts what products the user will be interested in from the user's past purchase history by using the internet history of e-commerce, recommends the next product to purchase, predicts the flow of future prices from the past price of stocks and transaction information in finance, and predicts how the winning rate will change when placing the next Go(바둑) stone in what position in the current situation of the Go paly. This is what artificial intelligence does well.



    그리고 주어진 데이터의 특징에 따라 데이터가 어느 범주(class 또는 category)에 속하는지 판단하는 것을 분류(classification)라고 합니다.


    Classification is the problem of identifying to which of a set of categories a new data belongs, based on the characteristics of the given data.


    주어진 데이터를 분류하기 위해서는 각 데이터마다 우리가 다룰 수 있는 형태로 표현할 수 있고, 각 범주와 얼마나 가까운지 또는 얼마나 유사한지를 계산하여 최종 판단해야 합니다. 이러한 척도를 데이터의 유사도라고 합니다.


    In order to classify a given data, each data must be expressed in a form that we can deal with and can be determined by calculating how close or how similar each data is to each category. These measures are called similarity of data.



    유사도를 설명하기 위해서 먼저 거리라는 개념을 소개하겠습니다. 그렇다면 서로 다른 두 데이터가 얼마나 유사한지 어떻게 평가할 수 있을까요?


    In order to explain the similarity, we may use a concept of distance. Then how we can evaluate the similarity between two different data?



    물론 데이터의 종류와 분석가의 관심사에 따라 ‘유사도를 재는 척도’는 다양할 수 있습니다. 가장 쉽게 생각해볼 수 있는 것은 두 데이터 사이의 거리를 계산하는 것입니다. 데이터는 순서쌍, 순서조 또는 벡터(vector)로 표현할 수 있으며, 예를 들어, 두 점 A, B 사이의 거리는 다음과 같이 dist(A, B)는 보통 과 사이의 차이의 제곱과 와 사이의 차이의 제곱에 루트를 씌워서 계산해 왔습니다.


    The measures of similarity do vary depending on the type of data and the analyst's interests. The easiest way is to calculate the distance between the two data. The distance between two points A and B is defined by


                         


    이를 Euclid 거리(distance)라고 부릅니다. 이 거리가 가까우면 두 데이터는 유사하다고 볼 수 있고, 거리가 멀면, 두 데이터가 서로 관련성이 떨어진다고 얘기할 수 있습니다. 그리고 어떤 데이터와 어떤 범주와의 거리가 가까우면, 이 데이터는 이 범주에 속해있다고 말할 수 있습니다.


    This is called the Euclidean distance. If this distance is close, the two data can be said to be similar, and if the distance is far, the two data are less relevant. And if the distance between any data and any category is close, it can be said that this data belongs to this category.


    예를 들어서, 아래 그림에서는 A와 B 사이의 distance가 A와 C 사이의 distance보다 작은 것을 즉, 와 는 가깝고, 와 는 멉니다. 거리 개념에서. 그러니까 B가 A에 가깝다는 것을 이해할 수 있고 또 같은 말이지만, 는 보다 에 더 가깝다는 것을 이해할 수 있습니다.


    For example, the figure below shows that the distance between A and B is smaller than the distance between A and C.



    그림입니다.
원본 그림의 이름: CLP00004e6c0004.bmp
원본 그림의 크기: 가로 785pixel, 세로 446pixel

                                             <---이 이미지는 화면에 있으니 자막에는 없어도 됨)


    그리고 또, 그 유사도를 거리를 이용해서 계산할 수도 있습니다. 다음은 거리 개념의 일반화된 개념으로 노름을 소개하겠습니다.


    Let me introduce 'the definition of Norm" that is a generalized concept of distance.


    벡터 에 대해 벡터 의 크기를 다음 |||| 로 나타내고, 그것을 의 노름(norm)이라고 부릅니다. 보통 실수의 크기를 할 때 절댓값을 사용했습니다. 이제 벡터 의 노름(norm)은 벡터의 크기를 나타내고, 다음과 같이 bar를 두 개 사용하여 표현합니다.

                            


    그래서 의 노름(norm)은 제곱근 의 제곱 플러스 의 제곱의 제곱근으로 정의합니다. 그리고 이것을 노름이라고 부릅니다.


    For a vector a (bold), the size of a is called a 'norm' and defined by



    즉 의 노름(norm)은 원점에서 점 A에 이르는 거리와 같습니다. 앞에서 설명한 거리하고 같은 개념입니다. 따라서 두 벡터 와 에 대해서 벡터의 노름은 두 점 A와 B 사이의 distance가 되며, 다음이 성립합니다. 즉 distance of A, B, 즉 dist(A, B) 는  의 노름은 으로 표현됩니다. 이 distance 정의는 2차원 벡터뿐만 아니라 3차원은 물론 고차원 벡터나 데이터에 대해서도 동일한 형태로 일반화될 수 있습니다.



                    


    That is, the norm of a is equal to the distance at the point A from the origin. It is the same concept as the distance described earlier. Therefore, for two vectors a and b, the norm of the vector a-b is the distance between the two points A and B. This definition can be generalized in the same form not only for 2-dimensional vectors, but also for 3-dimensional vectors as well as any high dimensional vectors or data.


    예를 들어, 3차원 벡터 , 에 대해서, 그에 대응하는 3차원 공간에 있는 두 점에 대해 다음이 성립합니다. 의 노름은 으로 표현할 수 있고, 이것은 원점에서 점 A에 이르는 거리로 이해할 수 있습니다. distance of A, B는 의 노름으로 로 쓸 수 있습니다.


                 (원점에서 점 에 이르는 거리)

              


    For example, for 3-dimensional vectors a, b and two corresponding points in 3-dimensional space, the following holds: The norm of a can be expressed as above, which can be understood as the distance at the point A from the origin. The distance of A, B can be written as the norm of a-b.


    그래서 지금 정의한 노름은 2차원 3차원 벡터는 물론 고차원 벡터나 다른 데이터들, 일반적인 벡터에 대해서도 적용할 수 있는, 크기를 잴 수 있는 척도가 됩니다. 노름(norm)입니다. 이 노름에 대해서 좀 더 자세히 알고 싶으시면 다음 주소에 들어가서 보시면 참고자료가 있습니다. http://matrix.skku.ac.kr/K-MOOC-LA/cla-week-1.html


    So, the norm defined now becomes a measurable scale that can be applied not only to 2-dim and 3-dim vectors, but also to high-dimensional vectors, other data, and general vectors. It is the concept of Norm.


    K-MOOC 선형대수학에 자세하게 n차원 공간 벡터에 대한 노름과 관련된 추가정보들을 확인하실 수가 있습니다. 여기서 노름의 식을 n차원 벡터에 대해서도 설명을 했고, 실제 노름을 계산할 때에도 코드를 이용하여, 위에 손으로 4차원 실공간   안에 있는 두 개의 점 사이의 distance도, 각각의 노름도 다음과 같이 dist(A, B) 로 계산할 수 있습니다. 앞에 소개한 주소 K-MOOC 선형대수학에서 노름에 대한 추가 정보를 확인하였습니다.


    You can check additional information related to the norm for the n-dimensional vector in detail in K-MOOC linear algebra. Here, the norm is also introduced for n-dimensional vectors, as well as the distance between two points in the 4-dimensional real space. If you would like to know more about the norm in English, go to the following link for reference. http://matrix.skku.ac.kr/LA/Ch-1/


    이어서 의 벡터 , 와 스칼라 에 대하여 다음이 성립합니다. 각 벡터의 노름은 언제나 0과 같거나 크다. 만일 의 노름이 0이라면 는 영벡터여야 된다. 벡터에 스칼라배 한 것의 노름은 k의 절댓값 곱하기 의 노름이 됩니다. ‘두 벡터들의 addition에 대한 노름은 각각의 노름의 addition 보다 언제나 같거나 작다’는 ‘삼각부등식’이 성립합니다.  (화면의 수식을 보세요)


    For a and b in , and scalar k, the followings hold.


        (1) ,    

        (2)

        (3)

                                                             <---화면에 있으니 자막에는 없어도 됨)




    예제로, 두 벡터 와 에 대해서 노름과 스칼라배한 벡터의 노름과 벡터 사이의  거리(distance)를 한번 확인해보면, 벡터 v를 정의하고, 벡터 w를 예제에서 주었듯이 ‘첫 번째 v의 노름, 그 다음에 2배의 w에 대한 노름, v 마이너스 w의 노름’을 다음과 같이 명령어로 (v – w).norm으로 구해줍니다. 실행을 시키시면 v의 노름은 이고, v와 w 사이의 distance 인 v-w의 노름은 임을 확인할 수 있습니다.

    (화면의 코드를 보세요)


    For example, let us check the followings for vectors v and w.

    1) norm of a vector

    2) norm of scalar multiple of a vector

    3) distance between two vectors


    This example can be practiced with the following code.


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    v = vector([1, 2, -4])  # generate a vector

    w = vector([-1, 3, 1])

    print("||v|| =", v.norm())  # find and print norm with v.norm()

    print("||2*w|| =", (2*w).norm())

    print("||v - w|| =", (v - w).norm())  # find distance거리 구하기--------------------------------------------------------------------



    이제 노름, 크기, 거리를 활용한 데이터의 유사도 비교에 대해서 설명하겠습니다. 벡터 사이에 distance를 확인할 수 있는 척도인 노름을 배웠으니까 그 개념을 이용하여 데이터 사이의 유사도를 구하는 방법입니다.


    Now, we will explain the similarity of data using norm (size, or distance). We have learned the norm, which is a measure of the distance between vectors, so we can use that concept to find the similarity between data.


    예제 2는 세 개의 데이터 A, B, C를 주고 세 벡터 중에서 B는 A와 C 중 어느 데이터에 더 가까운지를 판단하라는 문제입니다. 이것을 계산하기 위해서는, B와 A 사이의 distance 즉, (B-A)의 노름과 또 B와 C 사이의 distance인 (B-C) 벡터의 노름을 구해서 어느 것이 더 작은가를 찾으면 되는 겁니다. 그래서 A 벡터와 B 벡터, C 벡터를 정의하고, (a-b)의 노름, (b-c)의 노름을 각각 구해서 비교해보도록 하겠습니다.


    Example 2 can be practiced with the following code.


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    a = vector([0, 1, -7, 1])  # generate a vector

    b = vector([5, 2, -1, 3])

    c = vector([-2, 0, -4, 6])

    distAB = (a - b).norm()  # dist(A, B)

    distBC = (b - c).norm()  # dist(B, C)

    print("dist(A, B) =", distAB)

    print("dist(B, C) =", distBC)

    bool(distAB < distBC)  # True 이면 dist(A, B) < dist(B, C) 성립

    ---------------------------------------------------------------------


    Example 2 gives you three data A, B, and C. Determine B is closer to A or C. To do this, you just need to find the distance between B and A, that is, the norm of (B-A) and the norm of the vector (B-C), which is the distance between B and C, and find which one is smaller. So, let's define vectors A, B, and C. And then compare the norms of (a-b) and (b-c) respectively.



    (a-b)의 노름, 즉 dist(A, B)는 스퀘어 루트 66이고, (b-c)의 노름은 스퀘어 루트 71이니까. 이게 더 작습니다. 즉, B와 A 사이의 거리는 스퀘어 루트 66이고, B와 C 사이의 거리는 스퀘어 루트 71 이니까, B하고 A 사이의 거리가 더 가깝습니다. 즉 dist(A, B)가 dist(B, C)보다 더 작다는 것이 true인지 확인하라고 했으니까. True로 나타난 것입니다.


    Because the norm of (a-b), that is, dist(A, B), is the square root of 66, and the norm of (b-c) is the square root of 71, the distance between B and A is the square root of 66, and the distance between B and C is the square root of 71, so the distance between B and A is closer. The output shows that it is True.


    이와 같이 우리는 비교하고 판단하여 우리가 물어본 것이 True인지 아닌지를 항상 확인할 수 있습니다. 여기서 다른 벡터를 주고 dimension이 더 큰 5차원 벡터를 주고, 6차원 벡터, 7차원 벡터 등 어떤 벡터를 주더라도, 새로 생긴 이 세 벡터 사이에 distance를 다음과 같이 비교할 수가 있습니다. 여기서 각각 5차원 벡터 3개를 A, B, C로 주었습니다. 이 3개 벡터 중에서 A, B와 B, C 사이의 distance를 서로 비교해서 어느 것이 더 가까운지 즉, 이 경우는 A, B의 distance가 B, C의 distance보다 작은지를 물어봤을 때, 예스(Yes)라고 나왔듯이, 실제 숫자로 보더라도 A, B의 distance가 B, C의 distance보다 작은 것을 확인할 수 있습니다.


    In this way, we can always compare and determine if what we ask is true or not. Here, if you give another vector, for example, a 5-dim vector, a 6-dim vector, or a 7-dim vector, etc, we can always compare the distance between these three newly created vectors as follows: Here, three 5-dim vectors are given as A, B, and C. Of these three vectors, when you compare the distance between A, B and the distance between B, C and ask which one is closer, we will have answer in a second.


    일단, 우리 3주차 첫 번째 강의는 여기까지 하고 문제를, 여러분들을 위한 연습문제를 제시하겠습니다.


    Now we wrap up our first lecture in the 3rd week with one open problem. 


    거리 척도를 사용하여 유사도를 계산할 수 있는 데이터의 종류는 어떤 것들이 있는지, 즉, 노름을 이용해서 서로 더 가까운지 먼지, 어느 쪽 편에 더 가까운지를 판단할 수 있는 데이터들이 주위에 있으니까 알아보십시오. 또 노름척도로 유사도를 판단하기 용이하지 않은 데이터의 경우, 유사도를 판단하는데 사용이 가능한 다른 척도는 무엇이 있을지 생각해보시기 바랍니다. 이에 대한 답을 3주차 두 번째 강의에서 이어서 설명해드리도록 하겠습니다. 그리고 두 개의 7차원 벡터 사이의 거리(distance)를 직접 구해 보시길 바랍니다. 거기에 관련된 컴퓨터 코딩 도구(tool)를 만들어 놨습니다. 참고해서 거리척도 데이터에 활용하시면 되겠습니다.


    [Problem] What kinds of data can be used to calculate the similarity using a distance scale? That is, there are data that can be used to determine whether they are closer to each other by using the norm. Then try to find the distance between any two 7-dim vectors generated by yourself. You can refer to the relevant Sage codes and use it for distance scale data.   

      [Also you may think about a different data set that you can use other measure to determine similarity when the norm scale is not easy to determine similarity.]


    3주차 두 번째 강의에서는 <사잇각을 활용한 데이터의 비교>에 대해서 말씀드리도록 할 예정입니다. 수고하셨습니다.


    In the second lecture, we will continue to explain the answer to this problem and talk about the similarity of data using angles. Thank you.

    [3-2 pre]


    3주 2차시 ‘데이터의 분류’에서는 사잇각을 활용한 데이터의 비교, 코사인 유사도의 개념과 이를 계산하기 위하여 필요한 내적(inner product)과 사잇각(angle)을 소개합니다.


    [3-2] 'Classification of Data.' In this 2nd session of Week 3, we will introduce <the concept of cosine similarity> that uses angles between two vectors. In order to compute them, the concept of inner product is required.


    [3-2]  3주차 ‘데이터의 분류’, 2차시 강의에서는 <사잇각을 활용한 데이터의 비교>, <코사인 유사도>의 개념과 이를 계산하기 위하여 필요한 내적(inner product)과 사잇각(angle)을 소개합니다.


    [3-2] Lecture 3, 'Classification of Data.' In this 2nd session, we will introduce <the concept of cosine similarity> using angles between two vectors. We need the concept of inner product to compute them.


    *****


    안녕하십니까. K-MOOC [인공지능을 위한 기초수학 입문] 3주차 두 번째 강의, <사잇각을 활용한 데이터의 비교>로 시작하겠습니다.


    Now we start the lecture with 'Comparison of data using angles'.


    데이터의 유형과 분석가의 관심사에 따라 데이터 사이의 유사도를 재는 척도는 다양할 수 있다고 지난 시간에 말씀드렸습니다. 앞에서는 두 데이터 사이의 거리 (distance)를 노름(norm)을 이용하여 계산하여 유사도를 측정하는 척도로 삼는 문제에 관련된 예제를 보여드렸습니다.


    The measures of similarity between different data may vary, depending on the type of data and the analyst's interests. In the previous session, we have seen some examples, in which to measure the similarity, we computed the distance between two different data using norm


    그러나 데이터 분석가가 단지 데이터의 패턴에만 관심이 있는 경우, 즉, 방향에만 관심이 있는 경우는 거리척도는 적합하지 않을 수 있습니다. 예를 들어, 아래와 같이 좌표평면 상에 벡터로 표현된 두개의 데이터 와 는 위에서 보듯이 라는 벡터와 라는 벡터는 방향성은 유사하지만 거리는 는 짧고 는 굉장히 길어서, 거리척도로는 두 데이터가 관계가 없다고 판단될 수도 있습니다.


    However, distance scales may not be appropriate if data analysts are only interested in the pattern of data such as their directions. For example, two different data can be expressed as vectors on a xy-coordinate plane. Vectors a and b, as shown below, are similar in direction, but the length (norm) of a is short and the length (norm) of b is very long, so it may be determined that the two data are not well related on a distance measure.

                             그림입니다.
원본 그림의 이름: CLP00003aa00002.bmp
원본 그림의 크기: 가로 393pixel, 세로 265pixel

                                                             <---화면에 있으니 자막에는 없어도 됨)



    그래서 이런 경우에는 방향성을 척도로 삼으면 어떨까하고 생각해볼 수 있습니다.그래서 두 벡터 사이의 사잇각(angle, 각도)에 대한 수학적인 셋팅(setting)을 하는 방법을 소개시켜 드리겠습니다. 바로 ‘코사인 유사도의 개념’입니다. cosine similarity 라고 부릅니다.


    So, in this case, you can think about directionality as a measure. Now we introduce how to build up a mathematical setting for the concept of angle between two vectors. This can be used for the concept of cosine similarity on data.


    단지 데이터의 패턴, 방향에만 관심이 있는 경우에는 유사도를 어떻게 측정할까?


    How do you measure similarity if you are only interested in patterns or directions of data?


    이 경우 우리는 데이터의 패턴, 방향에만 관심이 있으므로, 두 데이터, 즉 두 벡터가 이루는 사잇각 로 유사도를 측정할 수 있을 것입니다. 예를 들어, 가 작으면, 즉 이 사잇각이 작으면, 데이터의 유사도가 높고, 가 크면, 즉 많이 벌어져 있으면, 데이터가 상관관계가 낮으며 유사도가 낮다고 판단할 수 있을 것입니다.


    In this case, we are only interested in the pattern and direction of the data, so we use the angle (theta, ) of two vectors to measure the similarity with the two data. For example, if the angle is small, then we say that two data are similar. If the angle is large, then there is a large gap and the data can be considered to be less related and so on.


    사잇각이 벡터의 내적(inner product)으로부터 정의되므로, 를 직접 계산하기 보다는 벡터의 내적을 이용하여 의 코사인 값으로 유사도를 측정할 수 있습니다. 이를 코사인 유사도(cosine similarity)라고 합니다.


    Since the angle between two vectors can be defined with the inner product, we can measure the similarity with the cosine value of by using the inner product. This concept will be called cosine similarity.


    코사인 유사도를 소개하기 위하여 먼저 벡터의 내적(inner product)을 정의하도록 하겠습니다.


    To learn the cosine similarity, we will first define the inner product of two vectors.


    두 벡터 와 사이의 내적(inner product)은 다음과 같이 정의됩니다.


    The inner product between two vectors a and b is defined as follows:


                              ,


    두 벡터의 첫 번째 component들을 곱해서 놓고 더하기 두 번째 component들의 곱을 놓아서 이 둘을 합친 것을 두 벡터의 내적(inner product)이라고 정의하겠습니다.


    To define the inner product of two vectors, multiply the corresponding components of two vectors, and then add these two values.  ( {rmbold a} cdot {rmbold b} = a_1 b_1 + a_2 b_2 )


    3차원 벡터인 경우에는 항이 하나 더 늘어나고, 4차원 벡터인 경우에는 4개 항을 더하는 식으로 표시될 것입니다.


    If you have three-dimensional vectors, then all you need is to add one more term. And if you have four-dimensional vectors, then just add two more terms.


    그리고 내적은 아래의 성질을 만족하는데, 대부분은 실수의 곱셈이 만족하는 여러분들이 그동안 많이 보아오시던 성질과 유사합니다. 실수 사이에서도 그랬듯이, 벡터와 벡터 자신의 내적은 의 노름의 제곱하고 같고, 그것은 언제나 0보다 같거나 크며, 내적이 0이려면 (노름이 0이어야 되니까, 필요충분조건으로) 가 0벡터여야 합니다.


    The inner product satisfies the following properties. Most of the properties are similar to those for the multiplication of the real number system. As in the real number system, inner product of vector a and a is same as the square of a's norm, which is always greater than or equal to zero, and a must be 0 vector if and only if the inner product is zero.


      ① ,


    두 번째로 내적은 교환법칙이 성립합니다. 즉, 와 의 내적은 와 의 내적과 같습니다.


    The inner product do satisfy the commutative law.


      ②    (교환법칙)


    그리고 분배법칙이 성립합니다. 즉 의 내적은 각각 의 내적하고 같으며, 모두 실수 연산에도  성립하는 성질입니다.


    and the distribution law also holds.

      ③   (분배법칙)


    그리고 의 스칼라배한 것과 의 내적은 와 의 스칼라배 한 내적과 같고, 또 와 의 내적을 구한 후에 스칼라배 해준 것과 같습니다.


    The inner product of scalar multiple of a and b is equal to the scalar multiple of an inner product of a and b.


      ④


    따라서 보통 실수에 성립하는 성질이 모두 내적에서도 유사하게 성립하는 것을 확인할 수 있습니다.


    It can be seen that all of the properties of real number hold similarly for the inner product.


    이 자세한 내용은 K-MOOC 선형대수학 1.2절, ‘내적과 직교’  http://matrix.skku.ac.kr/K-MOOC-LA/cla-week-1.html 내용에서 확인하실 수 있습니다. 추가내용을 궁금하신 분들은 확인하시기 바랍니다.


    You can read more on inner product in English on this link

         http://matrix.skku.ac.kr/LA/Ch-9/.


    이어서 사잇각을 정의하도록 하겠습니다.


    Now we define the angle between two vectors.


    벡터의 내적은 두 벡터가 이루는 사잇각과 관련이 있는데, 먼저 아래 그림에서 피타고라스 정리를 적용하면 다음을 쉽게 알 수 있습니다.


    In the figure below, it is easy to see that the inner product of vectors is related to the angle of two vectors from the Pythagorean theorem.

    그림입니다.
원본 그림의 이름: CLP000033bc0001.bmp
원본 그림의 크기: 가로 431pixel, 세로 350pixel

                                                             <---화면에 있으니 자막에는 없어도 됨)


    라는 벡터가 있고 라는 벡터가 있으면, 길이를 각각 라고 하고, 에서 로 가는 벡터는 -로 표시할 수 있습니다.

    이 그림에서 표현되는 관계는, 보통 고등학교 교과서에서 배웠듯이, 피타고라스 정리를 이용하면, , 이고 따라서 라고 표현할 수 있는데, 이 식을 벡터를 이용하여 다시 표현하면 다음과 같습니다. 로 쓸 수 있습니다. 의 길이가 벡터 의 노름이고, 의 길이가 벡터 의 노름이 되기 때문입니다. 그리고 의 길이는 바로 가 되어 위의 식이 성립합니다.


    Let a and b be the vectors, and ||a|| and ||b|| denote their norm, respectively. The vector from b to a is represented by a-b. From the Pythagorean theorem, we have , which is equal to . Simplifying this expression, we have

                .

    since the length of a (bold) is a norm of a (bold) and so on. And the length of c is  .


    이 성질로부터 내적의 성질을 활용해서 dist(A, B) 의 제곱은, 즉 는 로 표시되고, inner product를 풀어쓰면, 로 쓰여질 수 있고, 는 고, 는 이고, 와 가 가환이라 같으므로, 더하면 2배의 로 쓸 수 있습니다.


       



    From this property, we have

       




    이제 두 식을 비교하면,  가 됩니다. 즉, rm cos` theta = {{bold{a}} it  CDOT  rm {bold{b}} it ` rm} over {LEFT ∥ rm {bold{a}} it RIGHT ∥ LEFT ∥ rm {bold{b}} it RIGHT ∥ rm} 이 되겠습니다. 이 식은 라는 벡터와 라는 벡터가 주어지면, 언제나 성립하게 되고, 이 식을 만족하게 되는 가 항상 존재합니다. 우리는 이 를 와 벡터의 사잇각(angle)이라고 부르는 것입니다.


    Now we compare two equations, then we have

    So, .


    Given two non-zero vectors a and b, this expression always holds, and there exist that expression is satisfied. We call this , the <angle> of two vectors a and b.


    즉, 임의의 두 벡터가 있다면, 그 벡터가 ^2 에 있건, ^3 에 있건, ^4 에 있건, 또는 다른 임의의 벡터공간 V에 있더라도, 벡터의 내적을 정의한다면, 노름과의 관계가, 노름의 첫 번째 성질에서 나오듯이, 두 벡터가 주어지면 언제나 이 식을 만족하는 가 존재하게 되고, 그 를 두 벡터 와 의 사잇각(angle)이라고 얘기합니다.


    That is, although the vectors are in  ^2, or in  ^3, or in  ^4, or in any other arbitrary vector space V, if you define the inner product on it, then there exist an angle that satisfying this equation for any given two non-zero vectors. It is defined as the angle between the two vectors a and b.


    이제 사잇각을 이용하여 코사인 유사도를 계산해보겠습니다.


    Now we compute the cosine similarity using the concept of this angle.


    두 데이터 , 의 코사인 유사도는 다음과 같이 계산할 수 있습니다. 차원 공간 의 두 벡터 , 에 대해서도 동일한 공식이 성립합니다. 임의의 벡터가 있으면, 두 개의 벡터 사이에 이 식을 만족하는 가 항상 존재하므로 사잇각을 계산할 수 있습니다.


    The cosine similarity between the two data(vectors) a and b can be computed as follows. The same formula holds for any two vectors a and b in n-dimensional space . Since a satisfies the equation always does exist for any given two vectors, so the angle can be found.


    이 식을 다시 풀어쓰면, 이 벡터와 이 벡터의 내적(inner product)으로 표시할 수 있습니다. 또 점으로 표시되어 있어 dot product 라고도 부릅니다.


    Inner product is also called as a dot product because it is often marked with a dot.


    라는 벡터가 있고 벡터가 있으면, 두 벡터 사이에 이 식을 만족하는 , 이것을 우리가 사잇각이라고 하고, 이 값이 커지면 사잇각은 작아지고 유사도는 같은 방향이 되니까 유사도는 더 높아진다고 얘기할 수 있습니다. 여기서 , 은 원 데이터 , 의 크기가 어떤지 상관없이, 크기가 항상 인 단위벡터(unit vector)이므로, 코사인 유사도는 데이터의 크기와 데이터 사이의 거리는 무시하고 단지 데이터의 패턴 즉 방향만 고려하는 척도가 되는 것입니다.


                    



    Let a and b be two non-sero vectors, and  be an angle between them. As the value of increases, the angle becomes smaller and the direction is getting closer, which means the degree of similarity increases. Let's consider normalized vectors and which is always a unit vector regardless of the size of the original data a and b, so the cosine similarity is a measure that ignores the size of the data or the distance between the data, and only considers the direction which may means the pattern of the data.


    그림입니다.
원본 그림의 이름: CLP00003aa00001.bmp
원본 그림의 크기: 가로 852pixel, 세로 383pixel

                                                             <---화면에 있으니 자막에는 없어도 됨)


    이렇게 두 데이터의 코사인 유사도를 계산하여, 만일 코사인 값이 크면, 코사인 함수의 성질에 의해 사잇각은 작아지게 되고, 그래서 둘 사이의 유사도는, 둘 사이의 방향성은 더 같아지게 되는 것 입니다. 즉 유사도가 더 높아지는 겁니다. 이런 방식으로 데이터 사이의 패턴을 분석할 수 있습니다.


    If the cosine value is large, then the similarity is very close since the cosine function is decreasing on the interval [0, pi]. Therefore, if the cosine value is close to 1, then the direction between the two vectors is more identical and the similarity is getting closer. In this way, patterns between the data can be analyzed.


    첫째 시간에 배웠듯이 코사인 그래프는 다음과 같고, 값이 크다는 것, 즉 1에 가깝다는 것은 이 angle(각도)가 작다는 것이고, 값이 작다는 것은 이 angle이 멀리 90도 에 가깝게 벌어진다는 그런 의미입니다.


    As mentioned in the first lecture, the graph of cosine function is as follows. It shows that the value of the function is close to 1 means the angle between two vectors is small.


    그림입니다.
원본 그림의 이름: CLP000033bc0003.bmp
원본 그림의 크기: 가로 647pixel, 세로 395pixel

                                                             <---화면에 있으니 자막에는 없어도 됨)



    실제 예제에서 두 벡터 와 를 이렇게 주고, 두 벡터 사이의 내적을 구하고 코사인 유사도, 사잇각을 한번 계산해 보겠습니다. (화면의 코드를 보세요)


    In the following example, for the given two vectors v and w, we compute the inner product of the two vectors, the angle and the cosine similarity.



    먼저 두 벡터를 정의합니다. (1, 2, -4)와 (-1, 3, 1)을 정의하고, 와 의 내적(inner product)을 계산하고, 의 노름 즉 v.norm() 을 구해라. 마찬가지로 의 노름, 즉 w.norm() 을 구해라. (화면의 코드를 보세요)


     cosine similarity를 정의할 때, 값을 코사인 유사도라 했으니까, 그리고 angle은 값이 되는 값이니까, 그 는 arccos 값이 됩니다.

        

    모두 이렇게 정의를 한 다음에. ‘와 의 inner product를 계산하시오, 의 노름을 프린트하시오, 의 노름을 프린트하시오, cosine similarity를 소숫점이하 세자리수까지 구하시오, 또 angle 사잇각을 유효숫자 세 자리 수까지 구하십시오’하고, 명령을 줘서 실행을 시키면,


    Define two vectors are as below, and then compute the inner product, the cosine similarity and the angle in radian.


    Note that is a cosine similarity by definition and is a arccos of .  (See the code in the screen)


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    v = vector([1, 2, -4])  # 벡터 생성

    w = vector([-1, 3, 1])

    vw = v.inner_product(w)  # 두 벡터의 내적. 형식은 a.inner_product(b)

    vn = v.norm()

    wn = w.norm()

    cos_sim = vw/(vn*wn)  # 코사인 유사도

    ang = arccos(cos_sim)  # 사잇각(라디안)

    print("v.w =", vw)

    print("||v|| =", vn)

    print("||w|| =", wn)

    print("cos_similarity =", cos_sim.n(digits = 3))

    print("angle(rad) =", ang.n(digits = 3))

    # .n(digits = d)는 유효숫자 d자리의 근삿값

    ---------------------------------------------------------------------

    output :

    v.w = 1

    ||v|| = sqrt(21)

    ||w|| = sqrt(11)

    cos_similarity = 0.0658

    angle(rad) = 1.51


    다음과 같이 v.w는 1이 되고, unit vector가 되도록 만들었고, ||v|| = sqrt(21), ||w|| = sqrt(11)이고, cosin similarity는 0.0658이니까, 그리고 각도는 라디안으로 1.51이 되는 것을 확인할 수 있었습니다. 여기서 내적을 구했고, cosine similarity를 0.0658로 구했고, 사잇각은 0과 사이, 면 3.141592입니다. 0과 3.141592 사이에서 1.51쯤 되는, 중간쯤 되는. 각도로 얘기하면 로 얘기하면 한 45도 근처에 있게 되는 걸 확인할 수 있습니다. 물론 로 나타낼 수도 있습니다.


    This output means that v.w=1 and ||v|| = sqrt(21), ||w|| = sqrt(11).

    So, cosine similarity is 0.0658. Thus angle is 1.51 (rad) or 45 degree.




    그리고 코사인 유사도를 활용하여 ‘4차원 공간 에 있는 3개의 데이터’에 대해서 는 와 중 어느 데이터에 더 가까운지를 판단해보라고 하면, 똑같은 식으로 a, b, c 벡터를 정의해 주고, cosine similarity를 a, b 사이의 관계인 cosine similarity는 식 으로 구하고, 또 b, c의 cosine similarity를 로 구하면 됩니다. 그리고 이 두 개를 비교해 봅니다. (화면의 코드를 보세요)


    For the next Example, see the following code in the screen.


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    a = vector([0, 1, -7, 1])  # 벡터 생성

    b = vector([5, 2, 1, 3])

    c = vector([-2, 0, -4, 6])

    cos_simAB = a.inner_product(b)/(a.norm()*b.norm())  # 코사인 유사도

    cos_simBC = b.inner_product(c)/(b.norm()*c.norm())

    print("cos_sim(A, B) =", cos_simAB.n(digits = 3))

    print("cos_sim(B, C) =", cos_simBC.n(digits = 3))

    bool(cos_simAB < cos_simBC)

    ---------------------------------------------------------------------

                                                             <---화면에 있으니 자막에는 없어도 됨)


    Using the cosine similarity to determine whether the data B is closer to A or C in a 4-dimensional space. To do that, we only need to compute the cosine similarity.


     a, b가 b, c보다 그 방향성이 더 가까운가 계산해보면, 강의록에서 보셨듯이. A, B 사이의 cosine similarity는 –0.0448 이고,  B와 C 사이의 cosine similarity는 0.0856 이니까, A, B 사이의 각도가 B, C 사이의 각도보다 거의 절반 정도 되는 걸 확인할 수 있습니다.


    The cosine similarity of A and B is –0.0448, and the cosine similarity of B and C is 0.0856. So we can see that the angle between A and B is almost half of the angle between B and C.


    따라서 A, B 사이의 cosine similarity 0.04  는  B와 C 사이의 cosine similarity 인 0.08보다 더 작다는 것 즉, A, B가 더 가깝다는 것을, 또 B는 A에 더 가까운 데이터라는 것을 확인할 수 있습니다.


    Therefore, you can see that the cosine similarity between A and B is smaller than 0.08, which is the cosine similarity between B and C, that means that B is closer to A than C.


    여기서도 마찬가지로 4차원이 아니라 5차원 벡터여도, 또는 6차원 벡터여도, 같은 코드가 그대로 작용이 됩니다. 즉, 이게 2차원, 3차원, 4차원, 5차원이 아니라 100차원, 10,000차원 인 경우, 즉 여러분들이 빅데이터를 가지고 있더라도, 같은 식으로 벡터 사이에 cosine similarity를 이용하여 방향성이 얼마나 가까운 지를 바로 확인할 수 있다는 의미입니다.


    Likewise, the same code applies to vectors in five-dimensional vector space, or six-dimensional vector space. So what we did was not only for two or three-dimensional vectors, but also for 100 and 10,000-dimensional vectors. That means when we have a real big data, we still can use the same cosine similarity on vectors in the same way we did to determine how close the directions are.


    이제 몇 가지 연습문제를 제시하고 3주차 강의를 마치도록 하겠습니다.


    Now, let me present a couple of exercise problems, then finish today's lecture.


    먼저 어떤 데이터들이 cosine similarity를 사용하여 분석 가능할지를 찾아 생각해 보십시오.  (방향성이 중요한 데이터들, 즉 방향성이 같은지 다른지, 취향이 데모크라틱인지, 리퍼블릭인지. 다양한 성향들이 있을 수 있습니다.)


    First, think about what kind of data can be analyzed using cosine similarity.


    다음 문제로는, 두 개의 5차원 데이터들, 5차원 벡터들 사이의 내적과 사잇각을 구해보십시오. 여기에 cosine similarity tool을 만들어 두었으니까 다른 교재라든지 실생활에서 찾은 5차원 데이터들을 가지고 실제 사잇각과 내적을 구해보도록 하십시오. http://matrix.skku.ac.kr/math4AI-tools/cosine_similarity


    As for the following questions, try to find the inner product and angles between the two 5-dimensional vectors. We have created a cosine similarity tool here, so try to find the actual angle and inner product with other textbooks or five-dimensional data found in real life.

    http://matrix.skku.ac.kr/math4AI-tools/cosine_similarity 


    마지막으로 여러분들이 다양한 데이터를 확보할 수 있는 source를 알려드리겠습니다.

    Last, we tell you about the data source that you can freely obtain from the following links.


    국가에서 보유하고 있는 다양한 공공데이터들은 공공데이터포털, 정부데이터입니다.   https://www.data.go.kr/ 에서 확인할 수 있고, 공공데이터를 활용한 예를 한번 보시면 https://www.data.go.kr/tcs/puc/selectPublicUseCaseListView.do 정부의 공공데이터들이 활용된 다양한 사례들이 있으니까, 데이터를 사용해서 예측하고, 분류하고, 예측하고, 판단하고, 하는 사례들을 활용할 수 있습니다.


    Various public government data can be found from the public data portal.

    You can check them in https://www.data.go.kr/ and 

    https://www.data.go.kr/tcs/puc/selectPublicUseCaseListView.do .

     You can use these data to classify, judge and predict.


    이 외에도 국가 중점데이터들, 중요한 데이터들에 대한 정보나 데이터 자체도 이 자료에서 구할 수 있으니까

    https://www.data.go.kr/tcs/eds/selectCoreDataListView.do 를 참고해서 관련된 데이터 처리를 할 수 있는 데이터 과학자가 될 수 있는 인재가 되도록, 인공지능과 이런 코드들을 활용해서 여러분이 직접 데이터를 불러서 실제 활용하고 하는 사례들을 활용할 수 있도록 실습을 해보시기 바랍니다.


    To become a data scientist, learn how to use codes to practice using your own data.


    그리고 해본 내용들을 QnA에서 공유하시면 되겠습니다.


    And you can share what you've done in QnA and discuss.


    이제 3주차에 배울 내용들을 간단히 복습해 보도록 하지요.


    Now, let’s briefly review what we did in week 3.


    이번 3주 차에는 <데이터의 분류>에 대해서 학습을 하였습니다. 데이터를 분류하기 위해서 어떤 기준으로 분류하느냐, 데이터의 크기가 비슷한, 유사한 데이터들로 분류하느냐 또는 데이터 즉 벡터의 방향성이 같은 데이터들을 비슷한, 유사한 데이터로 분류하느냐를 생각해서, 데이터의 유사도를 확인할 수 있는 준거, 하나는 노름, 또 하나는 사잇각(angle)으로 구분하는 방법을 알려드렸고. 그 내용들을 데이터 유사도에 대한 개념, 노름, 노름을 이용한 유사도 비교, 사이각을 활용한 데이터의 비교에 대해서 소개하였고. 실제 계산을 하기 위해서 필요한 내적의 개념을 소개하고 사잇각을 계산하는 코드를 제공해서 아래 웹 주소에서 실습을 실제 해보면서 설명했습니다.


    In this week, we learned on data classification. In order to classify the data, we have to check the similarity of the data by length or direction etc. One can be done with vector norm, and the other can be done by computing the angle (or inner product).

    http://matrix.skku.ac.kr/math4AI-tools/distance_similarity/



     The various concepts of similarity were introduced, such as inner product, norm and angle. And we now have codes to compute them.

     


    사람과 사람사이의 유사도를, 거리로, A와 B 사이의 거리하고 A와 C 사이의 거리를 비교해서 ‘A하고 B가 더 친하다.’라고 얘기를 할 수 있었고, 그 것을 활용하기 위해서 노름 척도를 사용했습니다. A가 추구하는 방향과 B가 추구하는 방향이 angle 로 가깝다, 이 둘이 C보다 가깝다, ‘같은 성향이다.’라는 것을 cosine similarity를 이용해서 소개했습니다. 이때, 내적 개념을 소개했고, 이 내적 개념을 n차원 벡터에 대해서 또 일반적인 벡터들에 대해서도 활용된다는 것을 이해하였습니다.


    Comparing the distance between A and B and the distance between A and C, we could say 'A and B are more friendly,' and we used the concept of norm to make this conclusion. We introduced the concept of cosine similarity which can show 'the similar tendency'. In order to do so, we introduced the inner product for n-dimensional vectors. All these tools can be used for any general vectors and data.

     

     

    다음 4주차에는 <선형연립방정식과 연립방정식의 해집합>에 대해서 학습하도록 하겠습니다. 수고하셨습니다.


    In the next week, we will learn about <System of linear equations and the solution set of it>. Thank you.





    그림입니다.
원본 그림의 이름: Math4AI-0.jpg
원본 그림의 크기: 가로 494pixel, 세로 346pixel

     


    Week 4. Introductory Mathematics for Artificial Intelligence

     

                인공지능을 위한 기초수학 입문

     


    4.  선형연립방정식의 해집합 48

    4.1 선형연립방정식

    4.2 첨가행렬

    4.3 가우스 소거법

    4.4 연립방정식의 해집합


    [4-1 pre]


    안녕하세요. 4주차 ‘선형연립방정식의 해집합’에 대해서 학습하겠습니다.


    주어진 데이터로부터 적절한 모델을 찾는 문제는 선형화 과정을 거쳐서 선형연립방정식으로 표현됩니다. 이번 주에는 먼저 선형연립방정식에 관한 몇 가지 용어와 주요 성질에 대하여 살펴보고, 실제 주어진 선형연립방정식의 해와 해집합을 구하는 가우스 소거법에 대하여 학습합니다.


    [4-1 pre]

    Hello. In today’s lecture, we will study the ‘the solution set of a linear system of equations.’


    The problem of finding an appropriate model from the given data can expressed as a linear system of equations through a linearization process. We will learn some terminologies and main properties of the system of linear equations, and will study the Gaussian elimination method to find a solution and the solution set of a given system of linear equations.


    반갑습니다. 여러분. K-MOOC [인공지능을 위한 기초수학 입문] 4주차 강의를 시작하겠습니다.


    Welcome to '[K-MOOC] Introductory Mathematics for Artificial Intelligence' class.


    1주차, 2주차, 3주차 강의에서는 인공지능에 필요한 기본적인 배경 지식들을 학습하였습니다. 이제 4주차에는 인공지능에 필요한 수학 중 코어(core)인 선형대수학에 대한 내용을 소개하도록 하겠습니다.


    So far, we have covered fundamental mathematics for AI. In this fourth week, I will introduce the core of math for AI, 'Linear Algebra'.


    먼저 4주차에는 ‘선형연립방정식의 해집합’에 대해서 이야기 합니다.


    Let's start with 'The solution set of system of linear equations'.

    http://matrix.skku.ac.kr/math4ai-intro/W4/ 


    선형대수학은 모든 문제를 해결 할 수 있는 수학의 몇 안 되는 분야 중 하나로 알려져 있습니다. Alan Tucker는 “선형대수학이야말로 수학적 이론이 추구하는 모델이다.” 라고 얘기했습니다. 이 절에서는 수학에서 가장 중요한 ‘선형연립방정식의 해집합을 구하는 방법’을 학습하겠습니다.


    Linear algebra is known as a subject that can be used to solve the most of our problems. Alan Tucker even said, "Linear algebra is the model that mathematical theory pursues."  In the first section, I will give you a method to find the solution set of a given system of linear equations.


    1. System of linear equations   선형연립방정식


    다음과 같이 미지수 와 에 관하여 일차식(linear, 직선)으로 표현되는 방정식을 미지수 , 에 관한 일차방정식(linear equation 또는 선형방정식) 이라고 얘기합니다. 즉, 같이 직선(line)으로 표시할 수 있는 방정식(equation)이 선형방정식의 예라고 할 수 있습니다.


    A linear equation in two variables x and y is any equation that can be written in the form: (방정식)

    For example, 2x-3y=1 is a linear equation.


    이 식을 만족하는 순서쌍 를 좌표로 하는 모든 점을 좌표평면 위에 나타내면 그래프는 직선(line)이 됩니다. 위 방정식을 Sage 코드를 이용하여 직접 그려 보면, 다음과 같이 직선으로 나타납니다.

    The solutions of this linear equation, which are the ordered pairs (x, y) that satisfy this expression 2x-3y=1, form a line in a plane.


    실제 이 코드를 카피해서 주소 http://matrix.skku.ac.kr/KOFAC/ 에서 실습을 해 보시면 됩니다.


    You can copy and paste the following Sage codes into a cell in the following link http://matrix.skku.ac.kr/KOFAC/ , then evaluate it.


    다음과 같이 여러 미지수들에 관한 유한개의 선형방정식의 모임을 선형연립방정식 (System of Linear Equations)이라고 합니다. 선형방정식들이 여러 개 있는 거죠. 그래서 시스템(system) 이라고 합니다. 선형연립방정식(System of Linear Equations)입니다. 다음과 같이 일차식, 선형방정식이 여러 개가 같이 묶여있죠. 이게 선형연립방정식입니다.


    A set of finite linear equations for variables x, y is called ‘a system of linear equations.’ The examples are as follows:


    그리고 이 선형연립방정식의 해는 두 선형방정식을 동시에 만족하는 의 값과 의 값(또는 순서쌍 )을 말합니다. 한 개의 선형방정식은 좌표평면에서 하나의 직선을 나타내므로, 위의 선형연립방정식의 경우, 두 개의 직선을 그려보면, 두 직선의 교점을 나타내는 순서쌍 , 이 점이 바로 해가 됩니다.


    The solution of this system of linear equations is the value of x and y that satisfies the two linear equations at the same time. Since one linear equation in two variables x and y represents a straight line in the coordinate plane, the above system of linear equations mean two straight lines. Thus the solution becomes the intersection of the two straight lines.



    아래 그림에서 알 수 있듯이, 일반적으로 개의 미지수를 가진 선형연립방정식은 다음 중 단 하나만(one and only one)을 만족합니다. 즉, 선형연립방정식은 이 3가지 중에 단 하나만 만족합니다.


    As you can see in the figure below, the system of linear equations satisfies one and only one of the following three cases.


    선형연립방정식은 (1) 유일한 해를 갖거나, (2) 무수히 많은 해를 갖거나, (3) 해를 전혀 갖지 않아 해집합이 공집합이거나, 위의 3가지 중에 단 하나만 만족합니다. 선형연립방정식이 2개의 해를 갖는다, 3개의 해를 갖는다는 경우는 존재하지 않습니다.


     A given linear system satisfies one and only one of the following

    (1) a Unique solution

    (2) Infinitely many solutions

    (3) No solution


    예를 들어서, 미지수가 2개인 선형연립방정식의 예와 그림을 확인해 봅시다.


    For example, we now have 3 systems of linear equations with 2-variables.


            (1)         (2)        (3)

                                                             <---화면에 있으니 자막에는 없어도 됨)


    따라서, 1번의 예는 유일한 해를 갖고, 두 번째 예는 무수히 많은 해를 갖고 세 번째 예는 해가 존재하지 않습니다. 두 개의 라인이 있다고 하면, 서로 한 점에서 만나거나, 무수히 많은 점을 겹치거나 또는 평행하거나 해서 해가 존재하지 않는 경우 말고 다른 경우는 존재하지 않습니다.


    We can easily check that, (1) has a Unique solution (2) has infinitely many solutions (3) and has No solution as we see in the figure.


    실제로 2변수 함수일 때만이 아니라 3변수 함수, 4변수 함수가 있는 선형연립방정식(system of linear equation)도 마찬가지로 해가 유일하게 존재하거나 또는 무수히 많은 해를 갖거나, 이 두 경우가 아니라고 하면 마지막 경우는 해가 존재하지 않는 경우밖에 없습니다.


    Furthermore, the above properties are also hold for any system of linear equations with n variables. That means any system of linear equations either has a Unique solution, or has infinitely many solutions, or has No solution.


    모두 아주 일반적으로 성립되는 사실입니다. 그래서 해집합에 대해서 학습하기 위하여 첨가행렬에 대해서 먼저 얘기하겠습니다.


    Let's talk about the augmented matrix to learn how to find the solution set of a given system of linear equations.




    2. Augmented Matrix 첨가행렬


    선형연립방정식은 행렬을 이용하여 표현할 수 있습니다.

    다음과 같이 개의 미지수를 갖는 개의 선형방정식을 벡터를 이용하여 표현할 수 있습니다.


    Any system of linear equations can be expressed in the form of matrices. It is possible to express m linear equations with n variables using vectors as follows:


                                                             <---화면에 있으니 자막에는 없어도 됨)


    이것을 행렬에 벡터를 곱했더니, 벡터가 되는 이런 Ax=b 식으로 표현할 수 있습니다. 이렇게 정확히 같은걸 바로 확인할 수 있습니다. 이때 행렬 를 선형연립방정식 의 계수행렬(coefficient matrix)이라 하며, 에 를 붙여서 만든 행렬


                     [A VDOTS  rm {bold{b}} it ``]= {bmatrix{``a _{11} ``&``a _{12} ``&`` CDOTS ``&``a _{1n} ``&VDOTS &``b _{1} ``#a _{21}&a _{22}&CDOTS &a _{2n}&VDOTS &b _{2}#VDOTS &VDOTS &DDOTS &VDOTS &VDOTS &VDOTS #a _{m1}&a _{m2}&CDOTS &a _{mn}&VDOTS &b _{m}}}

     

                                                             <---화면에 있으니 자막에는 없어도 됨)


    을 선형연립방정식의 첨가행렬(augmented matrix)이라고 합니다.


    Using the matrix product, the above equations can be written as . The matrix is called the coefficient matrix of the given system of linear equations . And the matrix obtained from and is called the augmented matrix [A : b] of the equations.


    [예제 1]로 다음 선형연립방정식에서 첨가행렬을 구해보겠습니다.


    [Example 1] Find the augmented matrix of following system of linear equations.


                        .   

                                                             <---화면에 있으니 자막에는 없어도 됨)


    이걸 로 쓰면 행렬 는 행렬이 되고, 는 (9, 1, 0)이라는 열(column)벡터가 되겠습니다. 그래서 에다 를 첨가(augment)시킵니다. 그래서 첨가행렬(augmented matrix)을 만들어서 프린트하라고 하면, 다음과 같이


    Using Sage codes, you can find the augmented matrix.

    The augmented matrix of the above system of linear equations is


    [A : b] =

    [ 1  1  2 : 9]

    [ 2  4 –3 : 1]

    [ 3  6 –5 : 0]


    첨가행렬이 나타나게 됩니다. 이것을 실습을 해 볼까요. 복사해서 실습실에 한번 가봅니다.

    실습실 http://matrix.skku.ac.kr/KOFAC/ 에 들어가서 보시면 여러분들이 실제 프로그램을 실행했을 때, 여러 가지, 미분, 적분, 행렬계산 모두 다 계산할 수 있는 것을 확인할 수 있습니다. (화면의 코드를 보세요)


    You can copy and paste the following Sage codes into a cell in the following link http://matrix.skku.ac.kr/KOFAC/ , then evaluate it.


    가지고 온 코드를 실습실의 빈칸에 채우고, 첨가행렬을 구해볼까요.


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    A = matrix([[1, 1, 2], [2, 4, -3], [3, 6, -5]])  # 계수행렬 입력

    b = vector([9, 1, 0])  # 상수항 벡터

    print("[A : b] =")

    A.augment(b)  # 첨가행렬

    ---------------------------------------------------------------------

                                                             <---화면에 있으니 자막에는 없어도 됨)


    [A : b] =

    [ 1  1  2 : 9]

    [ 2  4 –3 : 1]

    [ 3  6 –5 : 0]    ■


    그러면 아까 확인한 첨가행렬과 똑같은 첨가행렬이 나오게 됩니다. 이때, 행렬에서 숫자를 바꾸면 그것에 맞는 첨가행렬이 만들어집니다. 이런 식으로 복사해서 숙제하실 때 활용하시면 되겠습니다.


    교재에 제공된 코드들은 3×3 행렬 뿐만 아니라 사이즈가 큰 7×7, 10×10 같은 행렬들도 같은 식으로 첨가행렬을 만드실 수 있습니다.


    The code in the textbook can be applied for finding augmented matrices of any system of linear equations. not only to the given 3x3 matrix, but also to any size of matrices such as 7x7 or 10x10 matrices.


    차의 정사각행렬 가 가역이고 invertable 행렬이고 가 벡터일 때, 라는 연립방정식은 유일한 해를 갖지요. 가역행렬이니까 가 존재하고, 를 양변에 곱해주면 는 항등행렬(identity Matrix)가 돼서 만 남게 되고, 이것은 이 됩니다. 그래서 라는 유일한 해를 갖게 됩니다.


    If a square matrix A is invertible, the linear system of equations has the unique solution .


    [예제 2] 역행렬을 이용하여 다음 연립방정식 의 해를 구해보겠습니다.

    [Example 2]  Find a solution of the following system of linear equations, using an inverse of A. (See the following codes)


                        


    이렇게 연립방정식이 주어지면 앞에서 실습한 방식으로 이 코드를 입력할 수 있습니다.  (화면의 코드를 보세요)


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    A = matrix(3, 3, [1, 2, 3, 2, 5, 3, 1, 0, 8])

    b = vector([1, 3, -1])

    print("x =", A.inverse()*b)  # Solve it with the inverse matrix,

    print()

    print("x =", A.solve_right(b))  # .solve_right( ) can do the same

    ---------------------------------------------------------------------

                                                             <---화면에 있으니 자막에는 없어도 됨)


     x= (-1, 1, 0)  : a solution ■


    따라서 주어진 선형연립방정식의 해는 , , 이다.    ■


    x가 이렇게 x= (-1, 1, 0)로 나온다는 의미는 벡터 x가 은 -1이고 는 1이고 는 0이라는 의미입니다.


    The output x= (-1, 1, 0) means a solution , , . ■


    행렬은 3×3 행렬 이고, 벡터는 (1, 3, -1)로 정의할 수 있습니다. 그 결과  x= (-1, 1, 0) 라는 유일한 해를 줍니다.

    실제로 행렬의 역행렬(inverse)를 구해보시면 가역행렬(invertable)임을 바로 확인해볼 수 있습니다. 실제 이런 계산도 명령어를 복사(copy)하셔서 실습해보시면 바로 구할 수 있습니다. http://matrix.skku.ac.kr/math4ai-intro/W4/


    You can find more codes related to finding an inverse matrix at the following link. http://matrix.skku.ac.kr/math4ai-intro/W4/



    [열린문제 12] 다른 교재의 선형 연립방정식의 해를 위의 명령어로 구해보시기 바랍니다. 어떤 것이어도 괜찮습니다. 사이즈가 큰 것도 상관없고 그 해를, 여기서 행렬을 집어넣고 b 벡터를 집어 넣어주면, 같은 명령어로 해(solution)를 구하실 수 있을 것입니다. 실제 그 예를 실습을 해보도록 하지요.

    http://matrix.skku.ac.kr/2018-album/LA-Sec-3-5-lab.html 

    이것은 제가 선형대수학을 가르치는 강의실 중에 일부인데, 선형연립방정식의 해집합과 행렬에 대해서 학습할 수 있는 자료입니다.


    [Open problem 12]  Find a solution of your own system of linear equations by using the code that you learned.


    여기에 보시면, 자명한 해를 갖는 경우, 또는 무수히 많은 해를 갖는 경우에 대해서 실제 실습해본 결과들, 첨가행렬을 만들고, 그리고 첨가행렬의 REF를 구해보고, 그 다음에 해가 존재하는지를 확인해보면, 위 경우에는 보시면 아시겠지만 0, 0, 0, 0도 해가 되고, 그 외에도 무수히 많은 해가 존재하는 것을 확인할 수 있습니다.


    In addition, as you can see in the above process, some system of linear equations can have infinitely many solutions. But, in this case, the code A.solve_right(b) only give a particular solution x_0 of Ax=b.


    그 다음 추가로, 선형연립방정식 가 무수히 많은 해를 가진다면, 먼저 A.solve_right(b) 명령어로 구한 특이해 x_0 와  다음 (수반동차연립방정식) 의 무수히 많은 해들의 집합 S, 를 더하여 전체 무수히 많은 해집합 x_0  + S 을 구하는 과정을 확인할 수 있습니다. 그 내용은 다음 시간에 소개해드리도록 하겠습니다.


    If the system of linear equations has infinitely many solutions, the solution set consists a particular solution x_0 of Ax=b obtained with the code A.solve_right(b) plus the solutions of Ax=0. Then the set x_0 + S becomes the whole solution set for Ax=b. We will introduce the details in the next class.


    4주차 1차시 강의를 마치고 다음 시간에는 RREF와 가우스 소거법(가우스-조르단 소거법)에 대해서 4장 3절에서 학습하도록 하겠습니다. 수고하셨습니다.


    In the next lecture, we will study RREF and Gaussian elimination methods to solve the system of linear equations in general. You will enjoy it. Thank you.


    [4-2 pre]


    반갑습니다. 이번 4주 2차시에서는 선형연립방정식이 해집합 중에서 가우스 소거법을 구체적으로 학습하고 이를 그것을 이용하여 연립방정식의 해와 해집합을 구하는 방법을 학습하겠습니다.




    [4-2 pre]


    Welcome! In the second session of Week 4, we learn the Gaussian elimination method for solving the linear system of equations in detail, and learn how to obtain the solution set of a given system of linear equations using Gaussian elimination.


    [4주차 2강] 반갑습니다. 4주차 첫 시간에는 우리가 선형연립방정식의 해집합에 대해서 공부했습니다. 먼저 선형방정식과 선형연립방정식을 소개했으며, 선형연립방정식의 해의 종류가 세 가지 경우 중에 한 가지 뿐이라는 것을 학습하였습니다. 그리고 첨가행렬 (augmented matrix)를 정의하여, 선형연립방정식을 행렬표현 로 바꾼 후, 첨가행렬을 이용하여 선형연립방정식의 해를 구하는 방법을 역행렬을 이용하여 학습하였습니다. 또, A.solve_right(b) 명령어를 이용하여 (특수)해를 구하는 방법도 실습을 했고, 그 두 방법 모두 같은 해를 찾아준다는 것을 학습하였습니다.


    [Week 4, Lecture 2] In the previous class, we learned the solution set of  a system of linear equations. We found that any system of linear equations satisfies one and only of one the following three cases: Unique solution,  Infinitely many solutions, or No solution. And we learned how to find the augmented matrix of a given system of linear equations. We used some codes to find its particular solution.


    오늘 4장 3절, 가우스 소거법을 학습하도록 하겠습니다.


    Today, we will cover Sec 4.3, Gaussian Elimination.

    http://matrix.skku.ac.kr/math4ai-intro/W4/ 


    일반적으로 선형연립방정식의 해집합을 구하는 가우스 소거법(가우스-조르단 소거법)에 대해서 살펴보겠습니다.


    Let's take a look on Gaussian elimination (Gauss-Jordan elimination), which is an algorithm for solving systems of linear equations.


    우리가 선형연립방정식을 풀 때, 다음과 같은 식 이 주어지면, 첨가행렬이 점점 단순화되면서, 여기에는 항등행렬(identity matrix)와 (2, -1)이라는 값을 확인했습니다. 즉, 이것이 의미하는 것은 , 이라는 식입니다.

    즉, 선형연립방정식을 풀 때, 첨가행렬을 점점 단순화시켜서 계수행렬 부분이 항등행렬이 되도록 연산하면, 답을 그대로 얻을 수 있습니다.


    When we solve a system of linear equations using Gaussian elimination, if the final form (of augmented matrix) on the left becomes the identity matrix, then the solution can be found without any difficulty.


    위의 예시에서 우리가 사용한 3가지 연산을 확인해보면, 두 식을 교환하거나, 한 식에 0 아닌 실수배를 해주거나 한 식에 0이 아닌 실수배를 하여 다른 식에 더해주는 그런 3가지 연산을 한 것입니다. 이 3가지 연산을 기본행연산(Elementary Row Operations)이라고 부릅니다.


    Three operations that were used in the above Gaussian elimination process are called ‘Elementary Row Operations.’


      (1) 두 식을 교환합니다.                                

      (2) 한 식에 0 아닌 실수를 곱합니다.

      (3) 한 식에 0 아닌 실수배를 하여 다른 식에 더합니다.


    (1) Exchange two equations.

    (2) Multiply an equation by a non-zero real number.

    (3) Add a non-zero multiple of an equation to another equation.


    위의 예시에서는 (2)번과 (3)번 기본행연산은 사용되었고 ‘두 식을 교환한다는 (1)번 기본행연산’은 사용되지 않았습니다. 그러나, 미지수가 많고 방정식의 개수가 많은 선형연립방정식의 경우 소거법을 체계적으로 다루기 위하여 식의 순서를 우선 정리하고 (2)번과 (3)번 기본행연산을 쓰는 것이 맞는 방법입니다.


    In the above example, the first type elementary row operation was not used. However, to systematically apply the elimination method to a system of linear equations, this operation is also involved, especially for the case of many variables and equations.


    기본행연산은 주어진 선형연립방정식을 단순한 선형 연립방정식으로 바꾸어줄 뿐 해집합은 전혀 바꾸지 않는다는 것이 키(key) 아이디어입니다. 또한 선형연립방정식에 행한 기본행연산은 첨가행렬에 행한 기본행연산과 일치합니다. 따라서 행렬에 대해서도 기본행연산을 정의할 수 있습니다.


    The important thing is that the Elementary Row Operations do not change the solution set. The idea of Elementary Row Operations for a system of linear equations can be used for the augmented matrix as well. Now we define a Elementary Row Operations for a matrix.


      (1) 행렬의 두 행을 서로 바꿉니다.

      (2) 행렬의 한 행에 이 아닌 실수를 곱합니다.

      (3) 행렬의 한 행에 실수배를 하여 다른 행에 더합니다.


    (1) Exchange the i-th row and the j-th row of A

    (2) Multiply the i-th row of A by a nonzero constant k.

    (3) Add the multiplication of the i-th row of A by k to the j-th row.


    선형연립방정식을 푸는 가우스 소거법은 위의 예시에서와 같이 첨가행렬을 간단하게 하는 것입니다. 즉 기본행연산을 사용하여 첨가행렬을 왼쪽에서 오른쪽의 형태로 바꾸어 계산하는 것입니다. 다음과 같은 계수행렬을 기본행연산(Elementary Row Operations)을 취하여 앞의 부분이 대각선 행렬, 가능하면 항등행렬(Identity Matrix)이 되도록 단순화 시킨 후, 그에 대응하는 연립방정식의 해를 구함으로써, 원래 선형연립방정식의 해를 구하는 것입니다.


    Gaussian elimination is the process of simplifying the augmented matrix as shown above. Make the coefficient matrix to a diagonal matrix and if possible, identity matrix, and then obtain the solution easily from the simplified form.


    이 오른쪽 형태의 행렬모양을 원래 행렬의 기약 행 사다리꼴(reduced row echelon form, RREF)이라고 합니다. 주어진 행렬의 RREF는 아래와 같은 명령어로 구할 수 있습니다.


    The right-hand side matrix is called in the Reduced Row Echelon Form (RREF). Using the following Sage code, we can easily find the RREF of a given matrix.


    우선 RREF에 대한 공식적인 정의를 확인하면, 행렬의 RREF는 다음과 같이 정의됩니다.


    A matrix is called a reduced row-echelon form if it meets all of the following conditions:


     ⓵ 성분이 모두 0인 행이 존재하면 그 행은 행렬의 맨 아래에 위치한다.

     ⓶ 각 행에서 처음으로 나타나는 0이 아닌 성분은 1이다. 이때 이 1을 그 행의 선행성분(leading entry, leading 1)이라고 한다.

     ⓷ 행과 행 모두에 선행성분이 존재하면 ()행의 선행성분은 행의 선행성분보다 오른쪽에 위치한다.

     ⓸ 선행성분(leading entry in row)을 포함하는 열의 선행성분 외의 성분은 모두 0이다.



    (1) If there is a row consisting of only 0's, it is placed on the bottom position.

    (2) The first nonzero entry appearing in each row is 1. This 1 is called a leading entry.

    (3) If there is a leading entry in both the i-th row and the (i+1) row, the leading entry in the (i+1)th row is placed on the right of the leading entry in the i-th row.

    (4) If a column contains the leading entry of some row, then all the other entries of that column are 0.



    예제 3에서 주어진 행렬의 RREF를 구해보지요.


    [예제 3] 다음 선형연립방정식의 해를 계산하고, 직접 계산한 결과와 코딩을 이용하여 계산한 결과를 비교하시오. 이 연립방정식을 소거법을 이용해서 풀어서 답을 구하면, , , 이렇게 구해지는 것을 손으로 확인할 수 있습니다.


    [Example 3] Find a solution of the given system of linear equations, using Gaussian eliminations. (See the code)


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    A = matrix([[2, -3, 1], [7, 1, -2], [1, 4, 3]])  # codefficient matrix

    b = vector([10, 1, -9])  # vector b

    A.augment(b).rref()  # RREF( [ : ] )

    -------------------------------------------------------------


    [    1     0     0   39/59]

    [    0     1     0 -162/59]

    [    0     0     1   26/59]   # RREF( [ : ] ) ■


    This means that the solution is  , , .


    The following code also gives us the same solution.


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    A.solve_right(b)

    ---------------------------------------------------------------------

    (39/59, -162/59, 26/59)  # solution x= 39/59, y= -162/59, z= 26/59   ■


    위의 결과로부터 를 쉽게 구했습니다. RREF를 구해서 답을 구할 수도 있고, 내부명령어를 이용해서 Ax=b를 풀기 위해, A.solve_right(b) 명령어를 주면, 다음과 같이 특수해(답) 39/59, -162/59, 26/59 라는 하나의 답을 정확히 얻을 수가 있습니다. 이 두 개의 명령어만 사용하면 바로 답을 구할 수 있다는 것입니다.


    So we easily found a solution x= 39/59, y= -162/59, z= 26/59 using the code  A.solve_right(b).


    다른 연립방정식에 대하여도, 마찬가지로 계수행렬 A 와 벡터 b 를 정의해서 A.solve_right(b)나 RREF 명령어를 이용하여 답을 다음과 같이 구할 수 있습니다. 이제 여러분은 연립방정식이 주어지면 (존재하는 해는) 어렵지 않게 구할 수 있게 되었습니다.


    You can use A.solve_right(b) for other systems of linear equations.



    [Finding a Solution set for a system of linear equations]

                                [연립방정식의 해집합]


    이제 4절에서는 연립방정식의 해집합에 대해서 학습하겠습니다. 해가 유일하게 존재할 때는 다음과 같이 바로 하나의 답을 구할 수가 있습니다. 그런데 만일 해가 무수히 많거나, 또 해가 존재하지 않을 경우에는 어떻게 다루는지를 학습하겠습니다.

    즉, 선형연립방정식이 다음과 같이 주어져서, 미지수가 5개, 변수가 5개, 식이 3개라고 하면, 이에 대한 계수행렬 A 를 정의하고, 그 다음 벡터 b 를 정의한 후에, 첨가행렬(augment matrix)을 만들어서 RREF를 구합니다. 그럼 다음과 같이 구해집니다.


    In the section 4, we will talk about the solution set of the given system of linear equations.

    Now see the following system of linear equations. Define the coefficient matrix, the vector, and make the augmented matrix. Then find its RREF as follows:


    [예제4]의 연립방정식은 변수가 5개고, 미지수는 3개입니다. 이것에 대응하는 선형연립방정식을 써주면,

      식이 됩니다.


    이 RREF가 의미하는 것은 바로 이것입니다.


    The system in [Example 4] has 5 variables and 3 equations. The RREF corresponds the following system of linear equations.


           .



    따라서 와 의 값이 정해지면 그에 따라 , , 의 값이 정해집니다. , (여기서 , 는 임의의 실수)를 대입하면 다음을 얻습니다.


    Let , where r, s can be any real numbers. Then we have

                             

                                                         


    이 해를 벡터로 쓰면 다음과 같습니다.


    This can be written as a vector form:


             .

                                                          


    이와 같이 가우스 소거법은 해집합을 구할 때, RREF를 구한 다음에 단순화시켜서 남는 변수, 여기서는 과 이고 자유변수(free variable)라고 합니다. 자유변수를 이용하여 나머지 변수들을 표현하고, 다음과 같이 ‘두 벡터의 일차결합으로 모든 해를 표시’하면 됩니다.


    Sage의 내부 명령어가 주는 해는 일반해 중에서, , 인 경우, 이 해 하나만 줍니다. 그래서 이 해를 주어진 선형연립방정식의 ‘특수해(Special solution)’라고 합니다.


    In Example 4, the variables , after obtaining the RREF are called free variables. The remaining variables x, y, z can be specified with those free variables. Once the values of r and s are specified, for example, r=0 and s=0, then the values of x, y, and z are also specified. This solution obtained from the specified values of r and s is called ‘a particular solution.’


     해집합 전체를 구하는 방법: 무수히 많은 해를 가지는 선형연립방정식의 경우는, 해집합을 찾는 일반적인 방법으로 먼저 Ax=b의 특수해 x_0 를 right_solve를 이용해서 구하고, 그 다음 Ax=0의 해집합을 구해서, 이 2개를 합친, x_0 + S  모양의 집합이 바로 전체 해집합을 이룹니다.


    How to get the whole set of solutions:

    1) Find a particular solution x_0 of Ax=b using the code 'solve_right.'

    2) Find a general solution set of Ax=0, say S

    3) The whole set of solutions is the set x_0 + S .


    이때, 동차선형연립방정식 Ax=0의 해집합 S를 구할 때, A의 right kernel을 구하라는 명령어를 사용합니다. 그러면 그 결과 벡터들의 일차결합의 모든 집합이 바로 Ax=0의 해집합 S 가 됩니다. 거기에 앞에서 구한 특수해 x_0 를 더해주면 전체 해집합 x_0 + S  을 구할 수가 있게 됩니다. 그 자세한 내용은 K-MOOC 선형대수학 4주차 강의 http://matrix.skku.ac.kr/K-MOOC-LA/cla-week-4.html 를 참고하시면 됩니다.


    A detailed method for obtaining the solution set for the homogeneous system of linear equations can be found in

    http://matrix.skku.ac.kr/LA/Ch-2/ .


    다시 요약하면, 이와 같이 A.solve_right(b) 하면 특수해 하나를 구해줍니다. 그 특수해를 구하고 아까, 명령어 A.right_kernel() 를 이용해서 Ax=0라는 연립방정식의 해집합을 구해서 여기에 그 집합을 더해주면 바로 Ax=b의 무수히 많은 해들을 다 담는 해집합 x_0 + S 을 구하게 됩니다.


    In summary:

    1)  Use the command A.solve_right(b) for ‘a particular solution’

    2) A.right_kernel() for the solution set of a homogeneous system of linear equations.


    다음으로 [예제5] 다음 연립방정식의 해를 구해보죠. 이 경우는 해가 없는 경우입니다.


    [Example 5] Find the solution set of the following system.


                         (i.e,      )

                                                 


    이렇게 변수는 3개밖에 없는데 식이 5개가 있는 경우, 이것을 손으로, 소거법으로 계산해보면, , , , 그리고 또 , 가 나오게 됩니다. 이것을 다시 확인해보면, A를 정의해주고, b를 정해준 다음에 첨가행렬 [A : b] 를 구해서 RREF [A : b] 를 구해보면 다음과 같은 RREF 모양을 줍니다. 이 RREF가 의미하는 바는, 이를 식으로 바꿔놓으면, 아까 우리가 단순화시킨 이 식하고 같아지게 됩니다. 여기서 보시면 , . 가 0인 것은 좋은데, 다음 식에서 좌변은 0에다 z를 곱하면 좌변은 0인데, 우변은 1이 되고 이는 0=1이라는 식이 돼서 모순이 됩니다. 이 식이 의미하는 바는 RREF에서 이 네 번째 행이 0=1이라서 모순이 생기는 겁니다. 이런 경우는 불가능합니다. 즉 0에다 실수 z를 곱했을 때 1이 되는 경우는 존재하지 않습니다. 즉, 네 번째 식은 모순이 되므로 그런 경우에는 위의 연립방정식은 해가 존재하지 않는 것입니다.


    From the RREF of the augmented matrix, we have 0*z=1 in the fourth equation. So this system has no solution.


     If we use code 'solve_right' instead of 'RREF', it gives us “has no solutions” as its answer.


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    A.solve_right(b)

    ---------------------------------------------------------------------

    ...

    ValueError: matrix equation has no solutions ■


    이는 바로 해가 존재하지 않는다는 것을 의미하고, 이 때는 해집합이 공집합이 되는 것입니다. 그래서 solve_right 명령어를 주어 A.solve_right(b)를 계산해주면, 위의 matrix equation은 has no solutions 이라는 답이 나옵니다. 같은 의미입니다.


    You can copy and paste the code 'A.solve_right(b)' into a cell to check that system has no solution.


    선형연립방정식의 첨가행렬을 기본행 연산(ERO)에 의해 REF(또는 RREF)로 변형하며, 값의 조건에 따라 해가 결정됩니다. (화면의 수식을 보세요)


    From the RREF of the augmented matrix of a given system of linear equations, we can easily determine if the system has a solution or no solution. For example, if you have a row that has all zeros except the last one d_{r+1} which is not zero, then the system has no solution.  (See the matrix on the screen).

    그림입니다.
원본 그림의 이름: K-001.jpg
원본 그림의 크기: 가로 1000pixel, 세로 520pixel

                                                             <---화면에 있으니 자막에는 없어도 됨)


    지금까지 선형연립방정식의 개념과 가우스 소거법에 관하여 살펴보았습니다. 인공지능에서, 주어진 데이터에 대한 적절한 모델을 찾는 문제는 선형연립방정식을 푸는 문제로 귀결되는데, 이 경우 일반적으로 해가 존재하지 않는 경우가 많습니다. 이를 해결하기 위하여 가능한 최선의 해, 최적해(optimal solution)인 <최소제곱해(least square solution)>을 찾는 문제가 되는데, 최소제곱해를 구하는 방법은 다음 절에서 자세히 살펴보도록 하겠습니다.


    We have studied the notion of a system of linear equations and how to find its solution. In artificial intelligence, the problem of finding an appropriate model for a given data results in a problem of solving a system of linear equations, in which case there are usually no solutions.  In the next class, we study how to find an optimal solution when the linear system does not have any solution.


    [열린문제 13] ‘주어진 선형연립방정식이 유일해를 갖는지, 무수히 많은 해를 갖는지, 해가 존재하지 않는지를 판단하는 것’은 첨가행렬 의 RREF를 구하여 이것만 자세히 보면 바로 판단이 가능한 이유, 즉 위에 설명한 내용이죠. 그것을 이해한데로 설명해 보시면 되겠습니다.


    [Open Problem 13] 'Explain why we can determine a given linear system of equations has either a unique solution or infinitely many solutions, or no solutions, just after we find the RREF[A:b].' Explain others what did you understand.


    이번 배운 내용에 대해서는 4주차 아래 주소에서 확인하실 수 있습니다.

    http://matrix.skku.ac.kr/K-MOOC-LA/cla-week-4.html  


    You can find more in http://matrix.skku.ac.kr/K-MOOC-LA/cla-week-4.html . 




     [Review]     [복습]


    우리가 여기서 배운 내용을 실습해보면, 4주차에는 선형연립방정식을 그리고, 이라는 그래프를 그리면, 이것이 축과 만나는 이 점이 근이 되겠죠. 선형연립방정식의 경우도 마찬가지로 그림을 두 식(equation)을 그려주면 유일한 해가 있는 경우, 무수히 많은 해가 있는 경우, 해가 존재하지 않는 경우를 구분할 수 있고, 연립방정식이 주어지면 행렬 형식 Ax=b 로 바꿔주고. 그 다음에 계수행렬을 만든 후에 첨가행렬을 만들고, 그 다음 역행렬을 곱해주거나 또는 solve_right 명령어를 이용해서 연립방정식을 풀면, 바로 답을 구할 수가 있고. 이 두 가지 방법이 주는 해가 같다는 것도 확인할 수가 있었습니다.


    오늘 배운 가우스 소거법 내용에 따라서, 선형연립방정식이 주어지면 계수행렬을 정의하고 벡터 b를 정의한 후에 첨가행렬의 RREF를 구해서 x, y, z을 바로 확인할 수 있습니다.


    그리고 선형연립방정식의 경우 해가 무수히 많은 경우는 RREF를 구해서, RREF 구한 식을 이용해서 해를 자유변수에 대해서 실수 r, s를 정의해서 전체 해집합을 구해줍니다. 해집합을 일반적으로 구할 때는, 앞서 설명했듯이 Ax=b에 특수해를 solve_right 명령어로 구하고 Ax=0의 해집합을 kernel 명령어를 이용해서 구해서 그것을 더해서 해집합을 구하면 됩니다.


    Now we review today’s class. We learned how to find the solution set of a system of linear equations:

    1) Find a particular solution x_0 of Ax=b using the code 'solve_right.'

    2) Find a general solution of Ax=0, say S

    3) The solution set is the set x_0 + S .


    그리고 해가 존재하지 않는 경우는 RREF를 구해서 그 RREF 의 모양을 보고 한쪽에는 계수행렬에는 다 0이 되는데, b벡터는 0이 아닌 경우에, 모순이 생기는 걸 확인해서 해가 존재하지 않으면 공집합을 갖는 것을 확인할 수 있습니다.


    We learned that from the RREF of the augmented matrix of a given system of linear equations, we can easily determine that the system has a solution or no solution. For example, if you have a row that has all zeros except the last one d_{r+1} which is not zero, then the system has no solution.


    이를 통해서 선형연립방정식이 주어지면 해, 특수해, 그리고 해집합을 구하는 방법을 학습했습니다. 이렇게 이번 4주차에는 선형연립방정식의 해집합에 대해서 학습하였습니다.


    그래서 4주차에는 선형연립방정식의 특수해 그리고 해집합을 구하는 방식을 선형연립방정식, 첨가행렬, 가우스 소거법, 연립방정식의 해집합에 따라서 학습을 하였고, 선형연립방정식의 해는 ‘유일한 해가 존재하거나, 무수히 많은 해가 존재하거나 또는 해를 갖지 않는다.’ 이 3가지 경우만 존재한다는 것을 학습하였습니다. 선형연립방정식은 Ax=b라는 행렬모양으로 바꾸어 쓴 후에 첨가행렬을 만들어서 기본행연산을 취하는 과정을 가우스 소거법(또는 가우스 조르단 소거법)이라고 부르며, 그것을 RREF 명령어를 줘서 바로 구할 수 있다는 것을 확인하고, (REF 또는)  RREF를 구해서 그것이 연립방정식의 해가 존재하지 않는 경우, 또는 해가 존재한다면 해가 유일하게 존재하는 경우와 해가 무수히 많이 존재하는 경우를 구분해서 학습했으며, 그리고 그것을 이용해서 선형연립방정식의 해집합을 구하는 과정을 학습하였습니다.


    In this class, we learned about Gaussian elimination and the RREF to find the solution set of system of linear equations. More details (in English) can be found in http://matrix.skku.ac.kr/LA/Ch-2/ .


    다음 시간에는 5주차 정사영과 최소제곱문제를 학습하겠습니다. 

    수고 많이 하셨습니다. 감사합니다.


    In the next week, we will cover the projection and the least square solution. Thank you.






    Week 5. Introductory Mathematics for Artificial Intelligence

     

                인공지능을 위한 기초수학 입문

     


    5. 정사영과 최소제곱문제 58

    5.1 최소제곱문제

    5.2 최소제곱문제의 의미

    5.3 정사영과 최소제곱해

    5.4 데이터에 적합한 곡선 찾기


    [5-1 pre]


    안녕하십니까, 인공지능을 위한 기초수학 입문 5주차. ‘정사영과 최소제곱문제’입니다.


    Welcome to Week 5 lectures of <Introductory Mathematics for AI>.

    Today, we will cover  <the Orthogonal Projection and the Least Square Problem>.


    이번 5주차 수업에서는 정사영(projection)와 최소제곱문제(least square problem)를 학습합니다.


    In this lecture, we will learn about <Projection and Least square problem>.


    지난 시간에 살펴본 바와 같이 주어진 데이터로부터 수학적 모델을 찾는 문제는 선형연립방정식 문제로 귀결됩니다. 그러나 측정에서 생긴 오차의 영향을 줄이기 위해 대개는 미지수의 개수보다 많은 개수의 데이터를 사용하기 때문에 방정식의 수가 미지수의 개수보다 많은 선형연립방정식을 만나게 됩니다.


    As mentioned in earlier, the problem of finding a mathematical model for a given data can be transformed to a problem of finding a system of linear equations.


    However, in order to reduce the effect of the error in the measurement, the number of equations may be more than the number of unknowns because the number of collected data is usually lot bigger than the number of specified variables/unknowns.


     이런 경우에는 일반적으로 선형연립방정식의 해가 존재하지 않는 경우가 많으므로, 가장 근사한 최적해(optimal solution)을 찾기 위하여 정사영 개념을 이용한 최소제곱문제(least square problem)로 바꾸어 최소제곱해(least square solution)를 구합니다.


    In this case, we may have to expect that this system of linear equations has no solution in general, so we have to find a best possible solution.  This can be done by solving a least square problem using the orthogonal projection and the least square solution.


    이번 정사영과 최소제곱문제는 4개의 절로 이루어져 있으며, 1절에서는 최소제곱문제, 2절에서는 최소제곱문제의 의미, 3절에서는 정사영과 최소제곱해, 그리고 4절에서는 데이터에 적합한 곡선을 찾는 문제(best curve fitting)를 학습하도록 하겠습니다.


    This lecture on <the orthogonal projection and the least square problem> consists of 4 sections, (1) the meaning of the least square problem in Section 1, (2) the least squares problem in Section 2, (3) the orthogonal projection and least squares solution in Section 3, and (4) the problem of finding a curve that best fits the data in Section 4 (best curve fitting).


    [5주차 1강]


    안녕하세요. 반갑습니다. K-MOOC [인공지능을 위한 기초수학 입문] 5주차입니다.

     

    이번 주에는 5절 <정사영과 최소제곱문제(least square problem)> 에 대해서 학습하도록 하겠습니다.


    [Week 5 Lecture 1]  In this fifth week, we will cover <Orthogonal projection and Least square problem>.


    최소제곱법은 데이터들의 패턴과 분포(behavior)를 잘 표현하는 근사직선이나 근사곡선을 구하는 아주 직관적이며 간단한 방법으로, 수치해석이나 회귀분석. 영상처리, 로봇 위치, 상태 분석 등에서 아주 다양하게 활용되고 있습니다. 이 절에서는 최소제곱법(least square method)이 실제 문제에 어떻게 활용되는지 확인하고 계산을 어떻게 하여 구하는지를 구체적 예를 통해서 확인하여, 실제 여러분들이 최소제곱해를 구할 수 있도록 해드리겠습니다.


    The least squares method is the most intuitive way to obtain a best fit curve that represents the patterns of the data. Let's find out how we can get the best possible solution with this interesting least square method.

    교안은 웹주소 http://matrix.skku.ac.kr/math4ai-intro/W5/ 에서 확인하실 수 있습니다. 이 주소에서 실습하시기를 권합니다.


    You can practice the contents at the following link:

     http://matrix.skku.ac.kr/math4ai-intro/W5/


    다음과 같이 와 에 관한 2차원 데이터가 주어져 있다고 합시다. 이 점들을 좌표평면에 나타내면 아래 왼쪽 그림과 같이 표현할 수 있습니다. 점이     로 표현됩니다.


    Consider 4 points in the xy-plane (2-dimensional data).

    Say, (0,1), (1,3), (2,4), (3,4).


    우리는 직관적으로 와 의 관계를 아래의 오른쪽 그림과 같이 직선 또는 파란색 점선으로 나타낼 수 있음을 알 수 있습니다. 이는 주어진 점들에 대응하는 근사한 직선 또는 곡선입니다. 이 데이터들은 측정에서 나온 결과입니다. 이 데이터들이 어떤  모양으로 변할지, 대개 이런 식으로 데이터가 진행될 것이라고 예측할 수 있는 근사 직선(line)을 구할 수가 있습니다. 이 직선(line) 중에 이 데이터를 가장 잘 표현하는 직선을 구하는 문제가 최소제곱문제이고, 그렇게 해서 찾은 직선이 최소제곱직선, (least square line 또는 best fit line) 이라고 합니다.


    Intuitively, it can be seen that the relationship between x and y is expressed by straight line or blue dotted lines as follows. The problem of obtaining a straight line that best represents the given data is called <Least Squares problem>. The straight line obtained is called <the least square line> or <the best fit line>.


    이제 최소제곱직선을 어떻게 찾는지 한번 생각해 보겠습니다.

    와 의 관계를 가장 잘 보여주는 일차함수 를 찾아봅시다.


    Let's think about the simplest case, that is, the problem of obtaining a straight line y=a+bx that best represents the given data.


    가장 이상적인 상황은 이 모든 데이터 에 대해서 식 가 그 4 점을 만족하면 됩니다. 즉, , 도 이 식을 만족하고, 도 이 식을 만족하고, 도 이 식을 만족하면 됩니다. 그런 계수 a, b만 찾으면 되는 것입니다. 따라서 다음과 같이 미지수가 , 인 선형연립방정식과 행렬표현을 얻게 됩니다.  (화면의 수식을 보세요)


    The most ideal situation is when all data are points of y=a+bx. In other words, it is the case of y_{i} = a + b x_{i} for all i. To find coefficients a and b, write this situation in a matrix form as follows.


               ,       

                                                             <---화면에 있으니 자막에는 없어도 됨)


    즉, 미지수는 2개, 식은 4개인 선형연립방정식을 행렬표현으로 바꾸면, 4×2 행렬에다 2×1 벡터를 곱해서 4×1 벡터가 되는 이 식 를 만족해야 됩니다.

     

    So we have a system of linear equations with 4x2 matrix A and a vector u.


     우리가 이미 알고 있듯이, 이 직선을 구하기 위해서는 두 개의 데이터에 대한 정보만 있으면 충분한데, 측정에서 생기는 오차의 영향을 줄이기 위해서 대개는 미지수의 개수보다 많은 개수의 데이터를 사용하게 되므로, 다음과 같이 방정식의 수가 미지수의 개수보다 많은 선형연립방정식이 생기게 됩니다.


    In general, to reduce the influence of errors in the observations, the number of data more than the number of unknowns are used.


    이런 경우, 일반적으로 선형연립방정식을 만족하는 해, 즉 인 는 존재하지 않거나 쉽게 찾지 못하는 경우가 대부분입니다. 대신 와 사이의 거리(distance)를 최소화 하는, 즉, Au와 y 사이의 오차가 최소가 되는 근사해, (u 햇 hat)을 찾는 문제가 되고, 이 을 찾는 문제가 바로 최소제곱문제(least square problem) 입니다. 즉 즉, dist( )를 최소화(minimize)하는 그런 를 찾는 문제라는 의미입니다. 이 은 를 직접 만족하지는 않지만, 이것을 만족하는 최적해(optimal solution), 이것이 바로 최소제곱해의 의미입니다.


    In this case, we do not expect has a solution u. So, we try to find an approximate that minimizes the distance between and y. This problem is called <least squares problem>. is called the optimal solution even though it does not really satisfy .


    이 최소제곱해가 주어진 4개의 식 모두를 만족하지는 않지만, 그런 식 중에서 가장 오차(error)가 작은, 즉 거리가 짧은 그런 직선(line)을 구하려고 합니다.


    We will try to obtain a straight line with the least error even though it does not pass all four points.


     [Meaning of the least square problem]  [ 최소제곱문제의 의미 ]


    이 최소제곱문제의 의미에 대한 추가 설명으로, 이 최소제곱문제는 다음과 같은 해석이 가능합니다. 즉 각 데이터 에 대하여 를 일차함수 에 대입하여 얻은 값을 ( 햇, hat)이라 합시다. 즉  를 집어넣으면 이 나옵니다. 와 똑같지 않을 수가 있습니다. 즉 선형연립방정식의 해가 존재하지 않는 경우에는 와 이 일치하지 않아서, 오차가 발생하는 경우입니다. 와 이 항상 같으면 바로 그것이 유일한 해가 되는 것입니다.


    Let be the value obtained by substituting into for each data . If the system of linear equations has no solution, it means that there exists an error for some since and  are not same. If and    are same for all , then the line is the unique solution.


    하지만 대개 모두가 같지는 않은 상황이 발생하므로 이를 해결하기 위한 차선책으로 각 데이터의 (제곱)오차 가 최소가 되는 , 를 구하는 것입니다.


    Because there are cases where is not all zero, we will try to find , that minimizes this error .


    주어진 모든 데이터에 대하여 오차들을 다 구해서 더하면 다음의 에러(Error) 식이 얻어집니다.


    Adding all the errors for all the given data gives the following error function .


      


    이 에러식 는 결국 Au와 y 사이의 거리(distance)의 제곱(square)하고 똑같아 집니다. 앞에 우리가 배운 노름과 내적(inner product)이 어떻게 에러(error)하고 관계되는지를 쉽게 확인하실 수 있습니다. 따라서 최소제곱문제는 아래와 같이 오차 를 최소화(minimize)하는 문제로, 이 문제의 최적해(optimal solution)가 바로 우리가 찾는 최소제곱해(least square solution)가 되는 것입니다.


    This error is eventually equal to the square of the distance between Au and y. It is easy to see how the norm and inner product are related to errors. The least squares problem is a problem that minimizes error . The optimal solution to this problem is the least square solution .


    그러면 이 최소제곱해(least square solution)를 어떻게 구하느냐, 이 때 바로 정사영(projection) 이라는 개념이 필요합니다. 이제 정사영(projection)에 대해서 알아보도록 하겠습니다.


    In order to find the least square solution, we need the concept of <projection>.


     [Projection and least square solution] [ 정사영과 최소제곱해 ]


    먼저 다음 식을 만족하는 해 를 찾는 문제를 생각합니다.

                   에 대해서 먼저 생각합니다.


    Let's consider the problem of finding t satisfy the following.


      Find  t such  is minimum.


    이 문제는 시작점이 같은 두 벡터 와 에 대하여, 를 포함하는 직선과 사이의 거리가 최소가 되게 하는 를 찾는 문제입니다.


    For vectors a and x, this is the problem of finding t that minimizes the distance between x and the straight line containing a.


     의 위로의 정사영(projection of onto )은 빛을 위에 비추면 그림자가 축에 닿는 바로 이 벡터(점)이 됩니다. 이 벡터 p 는 와 방향이 같고 크기만 다르니까, 에 상수배 즉 t배 한 것과 같아집니다. 즉 p는 에 상수배한 와 같아지는 것입니다.

     

    The projection of onto is the vector p on the a-axis (see the screen). This vector p is the same direction as a and only differs in size. Therefore, it is expressed as p = t a for some t.


      는 와 평행이므로 를 포함하는 직선 위에 놓이게 되고 [그림 2]에서 보듯이 실수 의 값에 대하여 는 실선으로 표시된 선분의 길이를 나타냅니다. 직관적으로 중 와의 거리가 가장 짧은 벡터는 의 끝점에서 를 포함하는 직선 위에 수선을 내려 생기는 벡터 임을 쉽게 알 수 있습니다. 그리고 이때의 값이 의 해가 됩니다. 벡터 를 위로의 의 정사영(projection)이라고 합니다. 실수 를 구하기 위해 임을 이용하면 다음을 얻습니다.  (화면의 수식을 보세요)


    We can see thatrepresents a distance between and as shown in [Figure 2]. Intuitively, the shortest distance can be obtained when and . Such a is a solution for and the vector is called the projection of onto .

    Since , this t can be obtained as follows:


                

                                        


    [예제 1] 은 , 일 때, 위로의 의 정사영을 구하는 문제입니다. 그렇다면, 이 식에 의해서  = 니까 = 만 구해서 벡터에 곱해주면 되는 것입니다. (화면의 코드를 보세요)


    [Example 1] Find a projection of y onto x, for , .


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    x = vector([2, -1, 3])           

    y = vector([4, -1, 2])

    yx = y.inner_product(x)  # inner product  of x and  y

    xx = x.inner_product(x)  # inner product  of x and  x

    p = yx/xx*x  # find projection

    print("p =", p)

    ---------------------------------------------------------------------

                                                        

    p = (15/7, -15/14, 45/14)    ■


    실제 정사영을 구하기 위해서 이 코드를 복사하여 실습을 해 보시고, 3차원 벡터에서도 정사영을 구하고, 4차원, 5차원, 7차원, 10차원 또는 100차원 벡터에서도 바로 같은 식으로 정사영을 구할 수가 있습니다.


    You can use the same code to get a projection for any finitely dimensional vectors.


    이 [표 1]의 최소제곱문제 는 정사영(projection)과 x 사이의 차를 최소화(minimize) 하는 그런 t 를 구하는 문제입니다. 이를 위해서 , 를 각각 의 첫 번째, 두 번째 열(column)벡터라고 하면, , 열(column)벡터들의 일차결합으로 풀어쓸 수 있습니다. 즉 다음과 같이 나타낼 수가 있습니다.


    We can solve the least squares problem in [Table 1], similar to the problem of determining t that minimizes the difference between the projection of x onto a and the vector x. For this purpose, let , be the first and second column vector of , respectively. Then, we have .


    은 바로, 열(column)벡터들의 일차결합, 열공간과 관계 됩니다.


    Hence, the least square problem is related to the column space of column vectors of A.


    이것이 바로 Ax가 만드는 치역(Range), 즉 이미지들의 집합이 되는 것입니다. 이런 평면이 되는 것입니다. 이것과 y의 거리(distance)가 가장 작아지는 벡터, 즉 정사영(projection)을 찾는 문제가 됩니다.


    This plane is a set of images given by Ax for x. The least square problem can be interpreted as a problem of finding a projection.


    이 그림에서 보듯이 이 문제는 ‘시작점이 같은 세 벡터 , , 에 대하여, 과 를 포함하는 평면과 벡터 사이의 거리(distance)가 최소가 되게 하는 , 를 구하는 문제라고 이해할 수 있습니다.


    As shown in this figure, it can be understood as 'the problem of obtaining a and b, for three vectors  , , with the same starting point, the minimal distance between the plane containing , and the vector .'


    그래서 는 실수 와 에 따라 과 를 포함하는 평면에 놓이게 되고, [그림 3]에서 보듯이 실수 와 의 값에 대하여 는 실선으로 표시된 선분의 길이를 나타냅니다. 직관적으로 중 와의 거리가 가장 짧은 벡터는 의 끝점에서 과 를 포함하는 평면 위에 수선을 내려 생기는 벡터 임을 쉽게 알 수 있습니다. 그리고 이때의 와 의 값이 의 해가 됩니다. 따라서 실수 와 를 구하기 위해 , 임을 이용하고 정리하면 이 됩니다.


    As shown in [Figure 3], for the values of a and b, is the length of the marked real line. It is easy to see that the vector with the shortest distance from to can be obtained by when and . Such an a and b is the solution of . The conditions and imply that .


     이를 다시 풀어쓰면, 로 풀어쓸 수 있고, 그 얘기는 가 가역행렬이기만 하면, 이 행렬의 역행렬(inverse)을 양변에 곱해서 은 바로 이 식 을 만족하게 하면 됩니다.


     ⇔


    If is invertible, then =.


    즉, A가 열(column)들이 일차독립이면, 가 항상 가역행렬이며, 그럼 위의 문제는 항상 만 구하면 그것이 답이 된다는 문제입니다.


    In other word, if all columns of A are linearly independent, then is always invertible, so  = can be easily obtained.


    [참고] 중요한 것은 (행렬식 성질에 의하여) 이 데이터 에서 이면 의 역행렬이 항상 존재하므로, 을 다음과 같이 쉽게 계산할 수 있다는 것입니다. [Vandermonde matrix의 행렬식 참조]


    [Note] If for a data set , then is always invertible.


     이 = 이 우리가 바로 찾고자 하는 최소제곱해가 되는 것이고 행렬의 곱과, 행렬의 전치행렬(transpose) 및 행렬의 역행렬(inverse), 그리고 연립방정식에 대한 지식만 있으면, 언제든지 최적해(optimal solution)  = 를 구할 수가 있다는 의미입니다.


    This vector  = is the least square solution that we wanted.


    그 개념을 이해해서 을 직접 구해보겠습니다.


    Let’s find the least square solution in the following example.


    [예제2] [표 1]의 선형연립방정식 의 최소제곱해를 구하면, 다음과 같이 행렬 A 와 u, y 로 풀어쓸 수 있고, 이 제곱해의 해는 다음과 같이 구할 수가 있습니다.


    [Example 2] Find a least square solution of in [Table 1].


    풀이. , , 에 대하여 의 최소제곱해(의 해)는 다음과 같이 구한다.


    Solution: 

               


    만약 데이터 크기(size)가 아주 크고, 변수가 많더라도, 같은 식으로 의 역행렬만 존재하는 게 확인된다면, 최소제곱해(least square solution)를 단 한 줄의 코드로 구할 수가 있는 것입니다. (화면의 코드를 보세요)


    Even though the data set is large, if the inverse of exists, then the least squares solution can be easily obtained by the following code.

     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    A = matrix([[1, 0], [1, 1], [1, 2], [1, 3]])

    y = vector([1, 3, 4, 4])

    print("u* =", (A.transpose()*A).inverse()*A.transpose()*y)

    ---------------------------------------------------------------------


    u* = (3/2, 1)   (즉, a=3/2, b=1)    ■


    따라서 우리가 얻은 최소제곱직선(least square line)은 이고, 통계학에서는 이런 최소제곱직선을 구하는 문제를 선형회귀(linear regression) 문제라고 합니다. 이게 우리가 찾은 최소제곱직선  입니다. 앞의 4개 점에 가장 근사한(best fit) 직선(line)을 구한 것입니다.


    So we have the best fit line  y= {3} over {2} + x .


    이상, 최소제곱문제에 대한 개념과 실제 찾는 구체적인 방법을 학습하였습니다.

    잠시 쉬었다가 이어서 5주 두 번째 강의 최소제곱문제 다음 부분을 실습하면서 확인하도록 하겠습니다. 수고하셨습니다.


    We now find the best fit line. We will continue this in our second lecture. Thank you.


    [5-2 pre]


    반갑습니다. 5주 2차시 ‘정사영과 최소제곱문제’에서는 정사영과 최소제곱해, 데이터에 적합한 곡선 찾기(best curve fitting)에 대해서 학습 합니다.


    Welcome! In the second session of Week 5 on ‘the Orthogonal Projection and the Least Square Problem’. Today we learn about the orthogonal projection, the least square solution, and a best curve fitting.


    주어진 데이터로부터 선형연립방정식을 만들고 그것을 행렬표현으로 바꿔서 가장 근사한 best possible solution을 찾아서 best line fitting 또는 best curve fitting 즉, 가장 데이터에 근사한 직선이나 곡선을 찾는 최소제곱문제(curve fitting)에 대해서 학습하도록 하겠습니다.


    We will build up a system of linear equations from the given data, and convert them into a matrix expression to find the best possible solution of it. Then we learn about the best line fitting and the best curve fitting, which is to find a straight line or curve that best fit the data.



    [5주차 2강]  반갑습니다. <일반인을 위한 K-MOOC. 인공지능을 위한 기초수학 입문> 5주차 두 번째 강의입니다.


    [Week 5 Lecture 2] Let’s start the second lecture of Week 5.


    오늘 <정사영과 최소제곱문제>를 마무리하도록 하겠습니다. 먼저 최소제곱문제는 너무나 중요한 문제이기 때문에, 다시 한 번 검토하면서, 설명하겠습니다.


    We will finalize this interesting section on <Projection and the least square problem>.


    네 개의 점이 있으면, 그 네 개의 점들을 가장 근사하게 표현할 수 있는 직선(line)을 구하는 문제입니다. 빨간 점들을 가장 근사하게 표현할 수 있는, 그 점들의 행동(behavior)을 가장 가깝게 표현할 수 있는 파란색 점선을 얘기하는 것입니다.


    When four points are given, it was our problem for finding a straight line that best fit these 4 points. We found this blue dotted line.

     

    그렇다면 이 식 y=a+bx 라는 식은 이 4개의 점을 다 만족해야 되는데, 그러려면, 이 연립방정식을 세울 수 있고, 그것을 행렬표현으로 바꾸면 Au=y로 풀어쓸 수 있고, 이 식이 실제, 그림에서도 봤듯이, 이 4점을 지나는(만족하는) 직선은 구할 수가 없으므로, Au와 y 사이에는 오차가 생기고, 그 오차(distance)를 최소화(minimize)하는 을 찾는 문제가 최소제곱문제입니다.


    The equation y = a + bx must satisfy all four given points. We can write this equation as a matrix form, Au=y. But as you can see in the .\PICture, there is no such straight line. Since there is an error between Au and y, our problem (the least squares problem) was finding that minimizes the error (distance between Au and y).



    최소제곱문제를 풀기 위해서, 이 차이(error)를 를 놓으면 바로 , Au와 y의 거리(distance)의 제곱이 바로 차이(error) 가 됩니다. 이 오차 함수를 최소화(minimize)하는 문제가 되는데, 이게 바로 전형적인 최적화(optimization) 문제가 됩니다.


    Let be the squared error function. Then we have =Our optimization problem is to minimize this error function.


    여기서 얻어진 값이 바로 우리가 찾고자 하는 최소제곱해입니다. 최소제곱해 문제를 풀기 위해서는 정사영(projection)에 대한 개념이 필요하고, 즉 벡터 a 위로의 x의 정사영, 즉 최단 거리인 이 점에서 p 라고 한다면, p는 a하고 방향이 같은 벡터 이므로, a의 상수배인 ta 로 표현할 수 있고, 이 p 와 x 와의 차이인 w, 즉, x-p = x-ta 또는 ta-x 인 이 벡터와 a가 직교(orthogonal)이므로 벡터 x-ta 와 벡터 a의 내적이 0이고, 따라서 와  = 를 바로 계산해 낼 수 있었습니다.


    The solution obtained is the least square solution. To solve least square problem, we need the concept of projection, this is the key feature.

    We took for example to illustrate the projection of x onto a = and how to find a solution . The important thing is that x-p = x - ta is orthogonal to the vector a.


    이 정사영(projection)을 구해보면, 실제 벡터 (2, -1, 3) 와 벡터 (4, -1, 2) 벡터의  내적을 구해서  = yx/xx =   을 구하면, 정사영(projection) = = (15/7, -15/14, 45/14) 인 것을 확인할 수 있습니다.


    For vectors  y=(2, -1, 3) and x=(4, -1, 2), the projection of y onto x is  = = (15/7, -15/14, 45/14).


      같은 방법으로 인 해를 최소제곱법으로 구하면, 에 (1, 3, 4, 4)를 곱하면 (1.5, 1)이 되는 것을 확인했고, 실제 그것을 계산해보시면 (1.5, 1) 로 나와서, y=1.5+x 인 것을 쉽게 확인할 수 있습니다.  즉 이 우리가 찾던 최소제곱직선인 것을 확인할 수 있습니다.


    Let’s find a solution for  . In the same manner, we can easily find that  y= {3} over {2} +x   is the least square line with (1.5, 1) = .


    [요약] 주어진 데이터로 연립방정식을 세우고,  Ax=b라는 문제 또는 Au=y라는 문제에서, 를 구하면, 이것이 바로 최소제곱해(optimal solution, least square solution) 가 되는 것입니다. 이제 여러분은 어떤 데이터를 주더라도, 대응하는 최소제곱해를 구할 수가 있게 되었습니다.


    [Summary] For a given data, construct a system of linear equations, deduce that Au=y, and then compute a least square solution by using = .



    아래 코드에서, 점이 더 많아도 되고, 달라도 됩니다. ‘만일 우리가 데이터 점을 100개를 주고 그 100개에 맞는 최소제곱직선(least square line)을 구해라’ 해도 바로 답을 구해주는 것입니다. 실제 그 과정을 실습 해보겠습니다. (화면의 코드를 보세요)


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    A = matrix([[1, 0], [1, 1], [1, 2], [1, 3]])

    y = vector([1, 3, 4, 4])

    print("u* =", (A.transpose()*A).inverse()*A.transpose()*y)

    ------------------------------------------------------------------

    Output: u* = (1.5, 1) = .



    [데이터에 적합한 곡선 찾기] [Find a suitable curve for data.]


    2장의 [5.4]절에서는 실제 주어진 데이터에 적합한 곡선을 찾는 문제를 소개해드리겠습니다. 직선을 찾는 문제보다 훨씬 더 중요한 문제입니다. 그 이유는 직선보다도 곡선이 데이터들의 행태(behavior)를 더 잘 표현해주기 때문입니다.


    In this section [5.4], we will consider a problem of finding a curve that best fit a given data. It is more interesting because curves express data better than the straight lines.



    앞서 학습한 방법을 활용하여 와 의 관계를 가장 잘 보여주는 (best fit) 이차 근사식 도 찾을 수가 있습니다. 즉 다음과 같이 미지수가 , , 인 선형연립방정식과 그에 대응하는 행렬표현을 얻게 됩니다.


    Like as before, we can find a quadratic approximation  that best represent the nature of data.


    같은 식으로. 만약 이전과 같은 점 4개가 일차함수가 아니라 이차함수를 만족한다면, 네 점이 식 를 만족해야 됩니다. 그러면 x=0일 때 y=1, x=1일 때 y=3, 등으로 식을 만족합니다. 그러니까 첫 식은  a+2b+4c 이고, 둘째 식은 x 등, 4개의 식 모두를  만족합니다. 이 식을 모아놓으면 미지수가 3개, 식은 4개인 선형연립방정식이 됩니다. 행렬로 표현하면 다음과 같습니다.


              

                             

                       

                                                             <---화면에 있으니 자막에는 없어도 됨)



    Since all 4 points should satisfy the quadratic equation , they can be expressed in a matrix form as follows.


     이때 원하는 (best possible) ‘u를 찾는 것이 이차식인 최소제곱곡선(curve)을 찾는 것’입니다. 이차함수인 최소제곱곡선(curve)를 구하기 위해서는 u=(, , ) 만 찾으면 됩니다. 가장 근사한 (best fit) 최소제곱곡선(least square curve)을 찾는 것입니다.


    At this point, finding a vector u=(a,b,c) is same as finding the least squares curve.


      이차함수를 구하기 위해서는 세 개의 데이터(즉, 세 점)에 대한 정보만 있으면 충분합니다. 그러나 측정에서 생기는 오차의 영향을 줄이기 위하여, 대개는 미지수의 개수보다 많은 개수의 데이터를 사용하므로, 방정식의 수가 미지수의 개수보다 많은 선형연립방정식이 생깁니다. 따라서 다음과 같은 최소제곱문제가 생깁니다.


    To obtain this quadratic function, only three points may be sufficient. However, when there are lots of points, the least squares problem also occurs.  


    이를 오차로 표현하면 다음과 같습니다.

     

    The following function E(u) represent an error.


                   

                          <---화면에 있으니 자막에는 없어도 됨)


    모든 조건이 같으므로, 같은 방법으로, 선형연립방정식 Au=y 에서 = 을 바로 얻을 수 있는 것입니다. 이것이 바로 우리가 원했던 의 계수 a, b, c 를 얻게 됩니다. (화면의 코드를 보세요)


    So we can find coefficients a, b and c by using .


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    A = matrix([[1, 0, 0], [1, 1, 1], [1, 2, 4], [1, 3, 9]])

    y = vector([1, 3, 4, 4])

    print("u* =", (A.transpose()*A).inverse()*A.transpose()*y)

    -------------------------------------------------------------------

    u* = (1, 5/2, -1/2)       (즉, a=1, b=5/2, c=-1/2)    ■



    따라서 u* = 는 다음과 같이 u* = (1, 5/2, -1/2)  입니다. 우리가 손으로 열심히 계산한 (1, 5/2, -1/2) 과 같은 답이 얻어집니다. 따라서 최소제곱곡선(least square curve) 을 구한 것입니다. 이것이 주어진 4점에 가장 근사한 (best fit) 이차곡선이 되는 것입니다. 같은 방법으로 주어진 4점에 가장 근사한 3차식도 구할 수 있습니다.


    The output is u* = (1, 5/2, -1/2), so the least square curve is where  a=1, b=5/2, c=-1/2. In a similar way, you can find a cubic approximation for any given 4-points as well.


    같은 방식으로, 예를 들어 점을 100개를 주고 100개의 점에 가장 근사한(best fit)  차수(degree)가 49인 다항식(polynomial)을 구하라고 한다면, 같은 식으로, 식을 이렇게 만들어주고, 그래서 행렬을 만들어주고, u* =를 구하여 계산하면 u* 의 성분들이 모두 나오고, 그 성분을 계수로 갖는 49차 다항식을 써주면 되는 것입니다.

    이와 같이 최소제곱해를 구하여 최소제곱직선 및 최소제곱곡선(best curve fitting, line 또는 curve)을 구하는 것입니다.


    Similarly, when 100 points are given, it is also possible to obtain the least square solution for a polynomial of degree 49. (Vandermonde added some idea on this problem too.) 


     여러분들이 ‘실험실에서 매 시간 측정한 데이터들을 x축과 y축으로 표현한 후, 앞으로 어떤 결과의 데이터가 나올지 적당한 차수(degree)를 주어 근사곡선(curve fitting)을 구하는 것’이 바로 여러분들이 오늘 배운 최소제곱법(least squares method)의 해법입니다. 이를 위하여 여러분들이 행렬의 전치행렬(transpose), 역행렬, 행렬곱, 행렬과 벡터의 곱을 배운 것입니다.


    To see the least squares method we learned today, we must know the notion of transpose and inverse of a given matrix as well as the product of matrices and vectors.


      이제 [열린문제 14] 를 풀어보시고 문의게시판에 답을 공유해주세요.


     앞서 구한 방법으로 평면의 6개의 점에 가장 근사한(best fit) 3차의 최소제곱곡선을 구할 수 있음에 대해서 논의해보시기 바랍니다.


     [Open problem 14] Find a least square curve of degree 3 for given 6-points in .



      최소제곱법 <실습실>은 이 주소 http://matrix.skku.ac.kr/2020-math4AI/LSS/ 에 주어져 있으며, 그래서 최소제곱법 문제를 직접 실습해 보시면서, 확인하실 수 있고, 실제 아까 여러분들이 보았듯이 계수가 3/2과 1 로 나왔습니다.


    그리고 최소제곱법 내용을 활용하여 part 3. 미적분에서 배울 경사하강법, 최적해를 구하는 경사하강법에 대해서도 예습해 두시면 도움이 될 것입니다.

    마지막으로 정사영에 대해서 더 알고 싶으시면, 우리가 만든 Geogebra 도구 https://www.geogebra.org/m/ewP9ybUP 을 활용하시면 됩니다.


    You can practice what we learned today in the following link:

    http://matrix.skku.ac.kr/2020-math4AI/LSS/ 


    벡터(데이터) x의 a 위로의 정사영, 이 정사영의 개념이 최소제곱곡선을 구하는데 결정적인 역할을 한다는 것, 여러분들이 실험실에서 만든 데이터들로부터 원하는 (best fit) 다항식으로 바꿔주어, 그 함수의 극대극소를 구하고, 면적을 구하여, 이익을 최대화하고, 손실을 최소화 하는 문제를 다룰 수 있게 된 것입니다.


    The concept of projection plays a crucial role in finding the least square solution. It will make us to deal with the problem of maximizing profit or minimizing cost after switching our data to a best fit polynomial.


    이와 같이 실험실의 (이산적인) 데이터와 (연속함수의) 미적분학을 연결시켜주는 브릿지 역할을 하는 것이 바로 최소제곱법이라는 것을 이해하시면 됩니다.


    The least squares method is a bridge that connects discrete data to knowledge of calculus that deals with continuous functions.


    5주차 강의가 인공지능 학습에 큰 주춧돌이 될 것입니다.


    [Review]


    수고하셨습니다. 그러면 간단히 5주차에 배운 내용을 복습하도록 하지요.


    이번 주 5주차에는 <정사영과 최소제곱문제>를 학습했습니다. 여기서는 우리가 측정한 데이터들로부터 데이터에 가장 근사한 다항식을 구하는 최소제곱문제를 학습했습니다. 이를 위해 정사영이라는 개념을 소개했습니다. 정사영이라는 개념을 사용해서  = 라는 개념을 그림을 통해 정확히 이해시켰고, t 를 구하는 방식을 이차원일 경우에 유도하고 이것이 3차원일 때, n차원 일 때도 똑같이 확장된다는 것을 보여줬습니다.

    데이터가 아무리 많더라도 그에 대응하는 최소제곱곡선(curve fitting하는 line 또는 curve)을 구할 수 있는 방법으로 최소제곱해(least square solution)를 배웠습니다. 방법은 행렬 A 를 갖는 연립방정식을 만든 후에, 그 A 에 대한 를 계산하여면 그것이 최소제곱해입니다.


    In this week, we have learned how to find a best fit least squares curve (of any needed degree) from the data.


    다음 시간에는 6주차, 행렬분해, QR분해, 특잇값 분해(SVD, singular value decomposition)에 대해서 학습하도록 하겠습니다. 수고 많으셨습니다. 감사합니다.


    In the next class, we will study <Matrix Decompositions> such as LU, QR Decomposition and SVD(singular value decomposition). Thank you.




         그림입니다.
원본 그림의 이름: AA.22593426.1.jpg
원본 그림의 크기: 가로 620pixel, 세로 572pixel

     


    Week 6. Introductory Mathematics for Artificial Intelligence

     

                인공지능을 위한 기초수학 입문

     


    6.  행렬분해 (특잇값 분해) 65

    6.1 LU 분해

    6.2 QR 분해

    6.3 SVD (특잇값 분해)

     - 과제 (열린문제)-  74


    [6-1 pre]


    안녕하십니까. 6주차에는 ‘행렬분해 (QR-분해와 특잇값 분해)’에 대해서 학습하도록 하겠습니다. 이번 6주차에는 주요한 행렬분해 기법 3가지에 대해서 학습합니다.


    Welcome. In Lecture 6, we will study some matrix decompositions (LU decomposition, QR-decomposition and Singular Value Decomposition).



    인공지능은 데이터를 기반으로 자료를 분류하거나 미래를 예측하는 등 가장 합리적인 의사결정을 해줍니다. 그러나 이를 실제로 구현하기 위해서는 프로그래밍 언어를 거쳐 코딩을 해야 합니다.


    Artificial intelligence makes the most reasonable decisions based on Data, such as <classifying data> or <predicting the future>. However, to actually implement this, we have to code through a programming language.


     즉 우리가 ‘컴퓨터가 이해하고 계산하는 방식을 알아야 합니다.’ 지금까지 학습한 행렬 연산을 실제로 구현하기 위해 주로 사용하는 행렬 분해 기법 3가지를 간단히 소개하도록 하겠습니다.


    In other words, we need to know how our computer understand and compute them. Let's briefly introduce three matrix decompositions, which we use for Matrix computations in real life.


    이번 주에는 먼저 LU분해와 QR분해를 학습하고, 마지막으로 특잇값분해 (Singular Value Decomposition, SVD)에 대해서 학습하겠습니다.


    First, we study LU decomposition and QR decomposition, and then we will study the Singular Value Decomposition (SVD).


    먼저 LU분해는 행렬 를 Lower triangular matrix와 Upper triangular matrix의 곱으로 표시해서 라는 방정식을 아주 쉬운 연립방정식으로 바꾸어, 대입하면서 해를 구하는 데 사용하는 방법입니다. 컴퓨터가 특히 잘하는 기능입니다.


    First, LU decomposition is used to solve a system of linear equations. If the coefficient matrix is decomposed as a product of a lower triangular matrix and an upper triangular matrix, then computers can solve the system of linear equations very quickly.


    연립방정식은 LU분해를 이용해서 풀고, 그 다음 배우는 QR분해는 를 직교행렬 Q와 삼각행렬 R의 곱으로 표시하는 방법으로, 이는 해가 존재하지 않을 경우에도 사이에 차이가 가장 최소화 되게 하는 optimal solution을 찾을 수 있게 해줍니다. 이를 이용하여 해가 존재하면 바로 그 유일한 해를 찾아주고 해가 무수히 많이 존재하면 그중에서 가장 최적의 조건을 만족하는 해를 구해주며, 해가 존재하지 않을 경우에도 optimal solution을 찾아줍니다.


    The QR decomposition is used to find an optimal solution to a least square problem. If a matrix is decomposed as a product of an orthogonal matrix Q and an upper triangular matrix R, then computers can solve the problem very quickly and accurately.


    이 QR분해는 단순히 연립방정식을 푸는 것을 도와줄 뿐만 아니라 고유값과 고유벡터를 찾는 것까지도 해 주는 아주 유용한 알고리즘으로, 컴퓨터에 생명을 불어 넣어준 수학 이론입니다.


    This QR decomposition is a very useful algorithm that not only helps to solve  least squares problems, but also finds eigenvalues and corresponding eigenvectors of a given square matrix. It gave a life to Computers to be used in real problems.


    마지막으로 특잇값 분해(SVD)는 이 중에서도 가장 중요한, 특히 인공지능에 사용되는 가장 중요한 행렬분해 기법으로, 행렬을 직교행렬 U와 일반화된 대각선 행렬 시그마, 그리고 직교행렬 V의 transpose를 곱한 행렬로 표시해 줍니다.


    The Singular Value Decomposition (SVD) is the most important matrix decomposition technique used in AI. Using SVD, a matrix is expressed as the product of an orthogonal matrix U, the generalized diagonal matrix Σ, and the transpose of an orthogonal matrix V.


    이 과정은 컴퓨터가 효과적으로 수행할 뿐만 아니라, 정사각행렬 외에도 직사각행렬(rectangular 행렬)인 m×n 크기의 임의 행렬 A에 대해서도 Singular Value Decomposition은 가능하기 때문에, 이를 이용하여 대상행렬의 차원을 축소(rank reduction)함으로써, 우리가 필요로 하는 해를 아주 빠른 시간에, 가장 근사한 해를 얻을 수 있게 해줍니다.


    The SVD is effectively performed by a computer, and the SVD of an arbitrary m×n matrix A (rectangular/square) can be always found. It can be used for a rank(dimension) reduction of a given matrix in AI. And it allows us to get the best possible solution that we need in a very short time.


    [6주차 1강] 안녕하십니까. <K-MOOC> <인공지능을 위한 기초수학 입문> 6주차입니다. 이번 주는 가장 중요한 3가지 행렬분해(Matrix decomposition)에 대해서 학습합니다. 그 중 인공지능에 필요한 가장 핵심적인 개념인 특잇값분해(singular value decomposition)가 포함됩니다.  


    [Week 6, Lec 1] In this sixth week, we will study some matrix decompositions, including SVD (singular value decomposition), which is a key concept for artificial intelligence study.


      인공지능은 데이터를 기반으로 자료를 분류하고, 분석하여, 미래를 예측하는 등 가장 합리적인 의사결정을 하도록 우리를 돕습니다. 우리도 마찬가지로 문제를 이해하고, 우리의 수학적 지식을 이용하여 적절한 알고리즘을 설계하여, 그것을 실제 프로그래밍 언어를 사용하여 코딩을 하는 과정을 거쳐 결과를 얻고 그 결과(output)를 활용하여 합리적인 의사결정을 하게 됩니다.

       이 과정에는 많은 계산이 필요하며, 우리는 계산을 위해 컴퓨터를 활용하기 위해서는  컴퓨터가 이해하고 계산하는 방식을 알고, 컴퓨터의 언어로 부탁을 할 수 있어야 합니다.


    Based on data, Artificial intelligence (AI) analyzes the situation, predicts the future, and makes a reasonable decision. In this process, a lots of computations are needed and some algorithms are used. Therefore, we should understand how do computers work and use a programming language to code our algorithms.


      지금까지는 선형대수학과 관련하여 행렬과 벡터의 연산, 선형연립방정식의 가우스 소거법, 최소제곱문제 등의 이론과 실제를 학습하였습니다. 컴퓨터는 위의 연산을 효과적으로 구현하기 위하여 행렬분해 기법을 이용합니다. 이 절에서는 3가지 중요한 행렬 분해를 간단히 소개합니다.


      6주차 강의 내용은 실습실 http://matrix.skku.ac.kr/math4ai-intro/W6/에 소개되어 있습니다. 우리는 이론 강의 후, 이 실습실을 활용하도록 하겠습니다.


    Let me introduce three important matrix decompositions in order to compute our problems effectively in AI.


    1. LU decomposition        LU 분해


      6장 1절에서는 LU분해를 학습합니다. 정사각행렬 를 다음과 같이 A=LU로 분해합니다. 여기서 L은 하삼각행렬(lower triangular matrix)이고 U는 상삼각행렬(upper triangular matrix)입니다. 하삼각행렬은 주대각선 성분의 위쪽의 성분이 모두 0인 행렬이고, 상삼각행렬은 주대각선 성분의 아래쪽 성분이 모두 0인 행렬입니다.


    [Definition] For a square matrix,

    Lower triangular matrix : If all entries above the main diagonal are zero.

    Upper triangular matrix : If all entries below the main diagonal are zero.


      이렇게 주어진 행렬 A를 하삼각행렬 과 상삼각행렬 의 곱 로 표현하는 것, 이것이 바로 LU 분해(LU decomposition)입니다. LU 분해는 선형연립방정식을 풀 때 많이 사용됩니다.


    LU-decomposition: For a square matrix A, LU decomposition refers to the factorization of A, (with proper row and/or column orderings or permutations), into two factors―a lower triangular matrix L and an upper triangular matrix U. (Say  A=LU.)


      계수행렬이 삼각행렬인 경우, 해를 구하기 위한 계산에서 역행렬이나 행렬식을 몰라도  대입 만으로 바로 를 구해갈 수 있기 때문에 계산량이 많이 적어집니다. 예를 들어서, b=Ax에서 A가 LU로 분해될 수 있다 그러면, b=LUx 로 쓸 수 있고 여기서 Ux를 y라고 놓으면 b=Ly 로 쓸 수 있습니다. 그러면 는 L 과 b를 알면서 를 구하는 문제가 되며, 여기서 대입하면서 y를 쉽게 구할 수 있고, 그러면 우리가 U와 y를 알면서, 에서 x를 찾는 문제로 바꿔집니다. 이것도 상삼각행렬(upper triangular matrix)에 대한 연립방정식 풀이 문제이기 때문에, 아주 쉽게, 역시 대입만 해서도 x 를 구할 수가 있습니다.


    If the coefficient matrix is a triangular matrix, the computation cost for solving the linear system of equations is reduced a lot even without knowing the inverse matrix. For example, if A is decomposed into LU in the system of linear equations Ax = b, it can be written as b=LUx, and if we set Ux as y, then it can be written as b=Ly. The original problem can be changed to a simple equation for finding x.


       즉, 로 분해 한 후 와 를 아주 쉽게 대입하면서 풀어서 우리가 원하는 를 만족하는 해 를 쉽게 구하면 되는 것입니다. 참고로 계수행렬이 삼각행렬인 선형연립방정식을 해결하는 데 필요한 계산량은 대략 밖에 되지 않고, 또 차의 정사각행렬 분해를 계산하는 데 필요한 계산은 의 계산량만 필요합니다. 이것은 연립방정식 Ax=b를 그냥 푸는데 필요한 대단히 큰 계승(factorial)개의 계산보다 매우 적은 계산이므로 아주 효과적인 연립방정식 해법이 될 수 있습니다.


    If the coefficient matrix A is any nxn triangular matrix, then the computation cost for solving  Ax=b is only about , which is a very small number of computation than that needed to solve any usual linear system of equations.


     요약하면 (또는 , 는 치환행렬 이고 )로 분해하여, 선형연립방정식 (또는 이고 )를 다음과 같이 해결할 수 있습니다.


    [단계 1] 를 로 쓰고 우선 를 만족하는 를 구한다. (전진대입법)

    [단계 2] [단계 1]에서 구한 를 이용하여 를 만족하는 를 구한다. (후진대입법)

     

    In short, find and

    [Step 1] Write as and find that satisfies .

    [Step 2] Use obtained in [Step 1] to find that satisfies .


    이 방법은 (행렬의 크기가 커지면) 역행렬이나 RREF를 이용하여 연립방정식을 푸는 것 보다, 연산의 수가 급격히 적어져서 컴퓨터가 더욱 빠르게 답을 구해주는 방법입니다.


    This method shows that if the size of the matrix is large, the solution set can be obtained much faster using the LU-decomposition than using RREF.


    실제 예제를 통해 살펴보면, 직관적으로 이해가 됩니다.


    Let's look at the next example.


    [예제 1] 행렬 를 다음과 같이 3×3 행렬이라고 할 때, 분해를 직접 구해봅니다. 실제 행렬 A가 주어졌을 때 (치환행렬 P와 하삼각행렬 L 및 상삼각행렬 U를 구하여  A=PLU로 로 분해하는)   과정을 아래와 같이 코딩으로 바로 할 수 있도록 만들었습니다.   (화면의 코드를 보세요)


    [Example 1] Let . Then find a LU decomposition of A.


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    A = matrix([[2, 6, 2], [-3, -8, 0], [4, 9, 2]])

    P, L, U = A.LU()    # LU-decomposition. A == P*L*U.

    print(P)             # permutation matrix

    print()

    print(L)             # lower triangular matrix

    print() 

    print(U)             # upper triangular matrix

    ---------------------------------------------------------------------

    [0 1 0]             # 치환행렬 P

    [0 0 1]

    [1 0 0]


    [   1    0    0]  # L은 하삼각행렬

    [ 1/2    1    0]

    [-3/4 -5/6    1]


    [  4   9   2]  # U는 상삼각행렬

    [  0 3/2   1]

    [  0   0 7/3]    ■

      


      실습실 http://matrix.skku.ac.kr/KOFAC/ 에서 명령어를 복사(copy)해 놓고, 클릭하시면 바로 답이 구해지는 것을 확인하실 수 있습니다.


    이제 여러분들은 다른 교재에서 행렬의 LU분해를 하라는 문제를 찾아서 크기(size)가 이것보다 아주 큰 행렬이어도, 행렬만 바꾼 후 같은 명령어를 실행하시면, 쉽게 LU분해를 할 수 있게 된 것입니다.


    For any large matrix, you can get its LU decomposition by using the same Sage code.


      [예제 2] 를 분해를 이용하여 해결해 보라는 문제입니다.  (화면의 코드를 보세요)


    [Example 2] Solve the  , using LU decomposition.


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    A = matrix([[2, 1, 1], [4, 1, 0], [-2, 2, 1]])     # Ax = y

    b = vector([1, -2, 7])     # PLUx = b

    P, L, U = A.LU()         # LU-분해. A == P*L*U

    y = L.solve_right(P.transpose()*b)  # Ly = P.transpose()*b, PLy=b

    x = U.solve_right(y)     # Ux = y

    print("x =", x)

    print("x =", A.solve_right(b))  # same as that obtained from built-in function.

    ---------------------------------------------------------------------

    Output:                                                   


    x = (-1, 2, 1)

    x = (-1, 2, 1)    ■


     위에서 보셨듯이, 실제 앞에서 배운 연립방정식 Ax=b를 풀어서 x를 구하라는 명령어로 구한 답과 답을 비교해보시면 이 둘이 모두 같다는 것을 확인할 수 있습니다.


    As you can see in the above, the answers from these two methods are the same.


       행렬의 크기가 작을 때는 이 두 답을 구하는 과정에서 계산 시간에 큰 차이가 나지 않지만, 행렬의 크기(size)가 5차, 10차, 100차, 1000차 행렬로 커지면 엄청난 시간 차이가 나오게 됩니다. 그것이 중요한 부분입니다.


    We do not see much of difference in computation time in these two methods when the size of the matrix is small, but you will see a huge time difference if the size of matrix is big.


      2차 세계대전 중 탄생한 컴퓨터에, 이 LU분해를 적용하여, 빠르게 답을 구하는 것을 보여준 현대식 컴퓨터의 개척자 <폰 노이만(John Von Neumann)>은 다른 학자는 물론 군인과 일반인들에게 <컴퓨터의 유용성>을 확신시켜준 계기를 제공하였습니다.


    <John Von Neumann> applied this LU decomposition algorithm to computers invented during World War II to win the battle.


      즉, 모든 문제는 선형화를 거쳐서, Ax=b라는 연립방정식 푸는 문제로 바꿀 수 있고, 그 답을 손으로 구하는 것보다 컴퓨터를 이용하는 것이 얼마나 더 유용한 지를 <LU분해>를 통하여 보여주면서, <아! 컴퓨터가 사람하고 할 수 있는 그 한계를 완전히 넘어서는 엄청난 일을 할 수 있다>는 확신을 주었습니다. 즉 이것이 컴퓨터에 생명력을 불어 넣어준, 그리고 모든 사람이 컴퓨터가 얼마나 중요한지, 그리고 앞으로 얼마나 더 엄청난 일을 하게 될 것인지를 확신시켜준 그러한 시작점이었습니다.


    All problems can be transformed into a linearized problem to be solved with the LU-decomposition, and we know it is essential to use a computer whenever it is needed. Do not try to do simple and routine computations by hand when any computer can do better.


    2. QR분해(QR Decomposition)


       행렬의 QR분해에 대해서 학습합니다. 앞에서 배운 최소제곱문제를 해결할 때, 사용하는 방법이 바로 <QR 분해>입니다.


    <QR-decomposition> can be used to solve the least squares problem.


    QR분해는 LU분해보다 더욱 중요합니다. 실제 우리가 연립방정식을 풀 때, LU분해보다 더 많이 쓰입니다. 또 최소제곱해를 구할 때 바로 쓸 수 있는 것이 이 QR분해입니다. 해가 존재하지 않을 경우도, 무수히 많이 존재할 경우도 우리가 최적해(optimal solution)를 찾게 하는 것이 바로 이 QR분해입니다.


    QR-decomposition is more important than LU-decomposition since QR-decomposition gives us an optimal solution even for the problems that has no solution or has infinitely many solutions.


        (m GEQ n)의 행렬 가 주어졌을 때, 정규직교벡터들을 열로 하는 행렬 와 크기의 (가역인) 상삼각행렬 의 곱으로 을 표현하는 것을 QR 분해(QR decomposition)라고 합니다. 여기서 정규직교벡터들이란 자신의 크기(노름)가 모두 1이면서 서로 다른 벡터들의 내적이 모두 0인 벡터들을 말합니다. 따라서 는 을 만족합니다.


    Let A be a matrix with . The QR-decomposition is a factorization of a matrix A into a product A = QR with an orthogonal matrix Q and an upper triangular matrix R. Any orthogonal matrix Q satisfy  , which means that all columns are orthonormal.


       를 QR로 분해를 할 수 있다면, 즉 이라면 우리가 앞에서 배운 최소제곱문제  를 다음과 같이 해결할 수 있습니다. 


    If we have , we can solve the least squares problem   as follows.


       앞서 이 최소제곱문제의 해이므로 아래 관계식이 성립합니다.


                             

                   

                  

                                      <---화면에 있으니 자막에는 없어도 됨)

    Because  is the solution of the above least squares problem, the following relationship holds.


       즉, 행렬이 QR분해만 되면, Q와 R을 이용해서 최소제곱해 = 를 바로 구할 수 있으니 연립방정식의 해를 구하는 문제를 획기적으로 짧은 시간에 풀 수 있게 만들어준 것입니다.


    If the matrix can be decomposed by QR, the least squares solution =   can be obtained immediately.


      수학자이며 물리학자인 <폰 노이만(John Von Neumann)> 도 지적하였듯이 ‘QR분해가 바로 사람이 경쟁할 수 없을 만큼 컴퓨터가  잘하는 일이고, 우리가 컴퓨터를 활용하게 해야 될 동기를 갖게 된 이유’이기도 합니다.


     <John Von Neumann> pointed out that 'computers do QR-decomposition better than human, and that is why we need to use computers.'



    [예제 3] 행렬 A가 3차 행렬  A= {bmatrix{``1``&``0``&``0``#1&1&0#1&1&1}} 로 주어졌을 때, A의 QR분해를 직접 해 봅시다.


    [Example 3] Find a QR-decomposition of .


     [참고] QR분해를 하는 과정에 Gram-Schmidt 정규직교화법을 활용합니다. 아래 코드는 Sage가 제공하는 gram_schmidt() 함수를 일부 수정하여 좀 더 직관적으로 그 과정을 확인하도록, gs_orth() 함수 코드를 만들어 QR 분해를 확인하였습니다. 우리가 실제 개발한 Python기반의 Sage 코드로, 우리가 만든 코드를 활용하여 위에서 설명한 QR 분해가 작동되는 과정을 구체적으로 보여줍니다. 여러분들도 앞으로 필요할 때에는 우리가 제공한 코드를 활용하여 수정해 가시면서 학습하시면 되겠습니다.  (화면의 코드를 보세요)


    QR분해를 짠 코드가 다음과 같습니다.


    To find a QR-decomposition, we use a Gram-Schmidt orthonormalization process. The followings are the Sage codes for the QR-decomposition.


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    # using gram_schmidt() function in Sage to define gs_orth() function

    def gs_orth(A):  # define a function

            

        m, n = A.nrows(), A.ncols()  # matrix size

        r = A.rank()  # compute rank of matrix

        

        if m < n:  # check the matrix size

            raise ValueError("The number of rows must be larger than the number of columns.")

            

        elif r < n:  # check if the matrix is full column rank

            

            raise ValueError("The matrix is not full column rank.")

        

        [G, mu] = A.transpose().gram_schmidt()   # use gram_schmidt

        # transpose of Q

        Q1 = matrix([G.row(i) / G.row(i).norm() for i in range(0, n)])

        R1 = Q1*A

        Q = simplify(Q1.transpose())  # matrix whose columns are orthonormal

        R = simplify(R1)              # upper triangular matrix

        

        return Q, R


    A = matrix([[1, 0, 0], [1, 1, 0], [1, 1, 1]])  # define a matrix

    Q, R = gs_orth(A)  # QR decomposition

    print("Q =")

    print(Q)

    print()

    print("R =")

    print(R)

    print()

    print("Q*R =")

    print(Q*R)

    ---------------------------------------------------------------------

    Q =

    [         1/3*sqrt(3) -1/3*sqrt(3)*sqrt(2)                    0]

    [         1/3*sqrt(3)  1/6*sqrt(3)*sqrt(2)         -1/2*sqrt(2)]

    [         1/3*sqrt(3)  1/6*sqrt(3)*sqrt(2)          1/2*sqrt(2)]


    R =

    [            sqrt(3)         2/3*sqrt(3)         1/3*sqrt(3)]

    [                  0 1/3*sqrt(3)*sqrt(2) 1/6*sqrt(3)*sqrt(2)]

    [                  0                   0         1/2*sqrt(2)]


    Q*R =

    [1 0 0]

    [1 1 0]

    [1 1 1]   # 임이 확인되었다. ■  (화면에 있으니 자막에는 없어도 됨)


    여러분은 이제 예제 4와 같은 방법으로 다른 교재에 나오는 선형연립방정식의 최소제곱문제를 바로 구할 수 있습니다. 저 코드를 그대로 복사(copy)해서, 여기에 행렬 A와 벡터 b만 대입해 주시면, 바로 여러분들이 원하는 최소제곱해(least square solution) 를 구하실 수 있습니다. 실제 이 과정을 문제를 풀면서 설명한 동영상이 주소 https://youtu.be/BC9qeR0JWis 에 있으니 필요시 확인하시면 됩니다.


    이상, 행렬분해 3가지 중에 LU분해와 QR분해에 대한 설명을 마치고, 다음 2차시에서 가장 중요한, 인공지능 수학에서 가장 중요한, 특잇값분해(Singular Value Decomposition)에 대해서 학습하도록 하겠습니다. 수고하셨습니다.


    In this first session, we have covered <LU-decomposition and QR-decomposition>.  We will continue to study the SVD (singular value decomposition) which is the most important to.\PIC in this class.


    [6-2 pre]


    6주 2차시 강의는 행렬분해 중 ‘특잇값 분해’입니다.


    In this second lecture of Week 6, we deal with 'Singular value Decomposition.'


    1차시에서는 LU분해와 QR분해에 대한 이론과 실제 및 실습을 해보았습니다.


    In the first lecture, we reviewed the theory of and did practice on LU decomposition and QR decomposition.


    이제 마지막으로 행렬분해 중 가장 중요한 특잇값 분해(SVD)에 대해서 이론뿐만 아니라 실제 계산을 수행하는 실습을 본격적으로 해보도록 하겠습니다.


    Now, let's discover the SVD, which is the most important matrix decomposition in AI.


    [6주차 2강]  http://matrix.skku.ac.kr/math4ai-intro/W6/


    안녕하세요. 반갑습니다. 6주차 2차시입니다.

     

    이번 시간에는 <행렬분해> 중 가장 중요한, 그리고 <인공지능을 위한 기초수학 입문> 중 가장 중요한 <특잇값 분해 (SVD, singular value decomposition)>에 대하여 살펴보겠습니다.


    [Week 6, Lec 2]  In this second session, we will study <SVD (singular value decomposition)>, which is the most important to.\PIC in the study of AI.


    만일 여러분들이 선형대수학을 배웠다면, 선형대수학의 가장 중요한 부분인 행렬의 대각화에 대해서 학습을 하셨을 겁니다. 그렇지만, 행렬의 대각화가 모든 행렬에 대해서 가능한 것이 아니고, 특히 정사각행렬이 아닌 경우에는, 행렬의 대각화를 얘기조차 할 수가 없었습니다. 그 한계를 벗어나서, 임의의 직사각형(rectangular) 행렬(m×n 행렬) 이라는 아주 일반적인 행렬에 대한 <행렬의 대각화>를 생각할 수 있느냐에 대한 답이 바로 SVD(singular value decomposition)입니다. 


    Matrix diagonalization is a big part of linear algebra. However, a matrix is not always diagonalizable. And if we have a non-square matrix, we can not even talk about matrix diagonalization. Matrix diagonalization can be considered as a special case of SVD. Furthermore, SVD can be applied for any size of rectangular matrix.


    SVD는 모든 크기의 (m×n) 행렬에 대하여 항상 가능합니다. 어떤 행렬 A가 주어져도 SVD는 구할 수 있습니다. 즉, 모든 행렬의 SVD는 존재한다, 이것이 가장 중요한 부분입니다.

     

    The SVD always exists for any sort of rectangular or square matrices. This is the key feature of SVD.




    크기가 인 행렬 에 대하여 의 직교행렬(orthogonal matrix) 와 의 직교행렬 및 의 (일반화된 대각선) 행렬 가 존재하여 다음을 만족합니다.  (화면의 수식을 보세요)


     The singular value decomposition of matrix is a matrix factorization in the form of , where is an orthogonal matrix, V is an orthogonal matrix, and is an rectangular (generalized) diagonal matrix with non-negative real numbers on the main diagonal.


              (화면에 있으니 자막에는 없어도 됨)

            

                                                             <---화면에 있으니 자막에는 없어도 됨)


    행렬 A가 어떤 크기(size)의 행렬이건 상관없이 직교행렬과, 일반화된 대각선 행렬과, 직교행렬의 전치행렬(transpose)로 분해될 수 있다는 것. 이것이 바로 SVD입니다.

    Regardless of the size of matrix, any matrix can be factorized into the product of an orthogonal matrix and a generalized diagonal matrix and the transpose of an orthogonal matrices. This is an SVD of a given matrix.


    이때,  직교행렬 U는 를 만족하고 V도 을 만족합니다. 행렬의 크기(size)는 각각 m×m 행렬과 n×n 행렬입니다. 그래서 (m×m) 행렬 곱하기 (m×n 행렬) 곱하기 (n×n 행렬)이 되어서 딱 맞아떨어지게 되는 것입니다.


    Note that U satisfies  , and V satisfies .


    이때, 은 주대각선성분이 모두 양수들이 단조감소()의 순서로 배열된 (k×k) 가역 대각선행렬이고, 나머지 성분들은 모두 0인 행렬입니다. 우리는 SVD에서 0 아닌 부터 까지의 주대각선 성분에 있는 0 보다 큰 이 숫자들 이것을 특잇값(singular value)이라고 부릅니다. 또 배열할 때, 이 특잇값을 크기 순서대로 배열합니다. 그리고 나머지 성분은 모두 0입니다. 이 특잇값이 매우 중요합니다.


      is a k×k invertible diagonal matrix whose main diagonal components are positive, and this component is arranged in descending order(). The other entries are all zero.

    We say these to are singular values of A.


    특잇값(singular value)은 여러분들이 이전에 배웠을 수 있는 고윳값의 일반화된 개념이라고 보시면 됩니다. 그리고 이 특잇값들은 모두 0보다 크고, 크기 순서대로 배열되어 있습니다. SVD에서는 위의 행렬 의 대각선 성분들을 행렬 의 특잇값이라 하고, 의 열들을 의 좌특이벡터(left singular vector), 의 열들을 의 우특이벡터(right singular vector)라고 합니다.

    즉, 에서 직교행렬 는 를 만족하고, 직교행렬 는 를 만족한다는 의미입니다.


    Singular values are generalized concepts of eigenvalues. As mentioned earlier, diagonal components of are called singular values of the matrix , and columns of are called the left singular vector of , and columns of are called the right singular vectors of .

    Hence, U satisfies and V satisfies .


    이 은 k×k 정사각행렬입니다. 의 양의 대각선성분들을 제외하고 (일반화된 대각선) 행렬 의 나머지 부분은 모두 0인 것입니다.


     is a kxk square diagonal matrix whose main diagonal entries are positive singular values of A.

     

    [예제 5] 다음 2×2 행렬 가 주어졌을 때, 특잇값 분해(singular value decomposition, SVD)를 구하라는 문제입니다. (화면의 코드를 보세요)


    [Example 5] Find a SVD of  A= {bmatrix{`` sqrt {3} ``&``2``#0&sqrt {3}}}

    .

     ------ http://matrix.skku.ac.kr/KOFAC/ -------------

    A = matrix([[sqrt(3), 2], [0, sqrt(3)]])  # define a matrix

    B = A.transpose()*A

    eig = B.eigenvalues()

    sv = [sqrt(i) for i in eig]  

    print(B.eigenvectors_right())

    ----------------------------------------------------------------

    [(9, [(1, sqrt(3))], 1), (1, [(1, -1/3*sqrt(3))], 1)]    ■


                                                            <---화면에 있으니 자막에는 없어도 됨)


     ---------- http://matrix.skku.ac.kr/KOFAC/ ---------------

    G = matrix([[1, sqrt(3)], [1, -1/3*sqrt(3)]])

    Vh = matrix([1/G.row(j).norm()*G.row(j) for j in range(0,2)])   # transpose V

    Vh = Vh.simplify()   

    print(Vh.transpose())

    print()

    U = matrix([A*Vh.row(j)/sv[j] for j in range(0,2)]).transpose()    # U

    print(U)              # column vectors in U are left singular vectors of A

    print()

    S = diagonal_matrix(sv)

    print(S)

    print()

    print(U*S*Vh)

    ----------------------------------------------------------------

    [        1/2 1/2*sqrt(3)]

    [1/2*sqrt(3)        -1/2]        # V


    [ 1/2*sqrt(3)        1/2]

    [        1/2 –1/2*sqrt(3)]      # U


    [3 0]

    [0 1]                          # S


    [sqrt(3)       2]

    [      0 sqrt(3)]             # ■

                                                            <---화면에 있으니 자막에는 없어도 됨)


    위의 코드에 의해서 주어진 행렬의 직교행렬 V를 구했고, 그 다음 직교행렬 U를 구했고, 그 다음에 대각선, 일반화된 대각선 행렬 를 구했습니다. 이 셋을 곱하여 을 구해 보면, 우리가 원래 가지고 있었던 행렬 가 되는 것을 확인할 수 있습니다.


    With the code above, we have an SVD of A. And we can check that is same as the original matrix A.


    이렇게 위의 명령어를 이용해서 여러분들이 다른 행렬을 제공하면, 그리고 그에 맞게 행렬 크기(size) 코드를 수정해 주시면, 이렇게 코드를 조금만 손보시면 어떤 행렬이 주어지던 간에 SVD(singular value decomposition)를 구할 수 있다는 것을 확인하실 수 있습니다.


    If we modify a little on the code above, we can have SVD for most of the given matrices.


    이렇게 구한 직교행렬과 대각선행렬, 직교행렬의 전치행렬(transpose)을 곱하면 원래 행렬이 되는 것까지도 확인할 수 있었습니다.


    You can check that become the original matrix A.


    위의 예에서 보여준 수학적 알고리즘을 하나의 명령어로 간단하게 만든 아래의 코드로 여러분은 더 쉽게 특잇값 분해를 할 수 있습니다. 또 SVD(singular value decomposition)의 명령어를 사용 할 때에는 여러분들이 행렬자체는 간단한 정수행렬을 가지고 있더라도,  와 에 대한 고윳값과 고유벡터를 구해야 하므로, 그 중에는 유리수나 무리수, 복소수 등이 나올 수 있고, 자연수나 정수 그리고 유리수를 벗어나는 그런 행렬들이 나올 수 있으므로, 우리는 지금까지 연산을 대개 실수 행렬에 제한했었지만, 이제부터는 실수 안에서는 모든 연산이 가능하도록 실수체(Real Field)를 의미하는 RDF에서 행렬을 정의해주겠습니다.

     

    In the process of obtaining an SVD, the components of the matrix may not be an integer or a natural number. Therefore, we will define a matrix whose components are in RDF, which means the real double field, so that all operations are possible.


    그래서 [예제 6번]에서는 행렬 의 특잇값 분해를 실수체에서 구하고 임을 확인하겠습니다. 그러면 행렬 A를 실수체 상에서 정의하고, 그다음 USV를 A의 SVD(singular value decomposition)하라는 이 한 줄, 명령어를 가지고 USV를 구해서 U를 프린트하고, 를 프린트하고, 또 V를 프린트하라. 그리고 나서 실제 해서, A와 같은지를 확인해봐 이렇게 명령어를 줬습니다.


    So, let's practice in [Example 6] with the following Sage codes.


     ---------- http://matrix.skku.ac.kr/KOFAC/ -------------

    A = matrix(RDF, [[sqrt(3), 2], [0, sqrt(3)]])  # define a matrix

    U, S, V = A.SVD()  # SVD.  A = U*S*V'

    print("U = ")

    print(U)

    print()

    print("S = ")

    print(S)

    print()

    print("V = ")

    print(V)

    print()

    print("A = USV^T = ")

    print(U*S*V.transpose())

    -------------------------------------------------------------

    U =

    [ 0.8660254037844387 -0.5000000000000001]

    [ 0.5000000000000001  0.8660254037844387]


    S =

    [               3.0                0.0]

    [               0.0 1.0000000000000002]


    V =

    [                0.5 -0.8660254037844387]

    [ 0.8660254037844387                 0.5]


    A = USV^T =

    [     1.7320508075688776                     2.0]

    [-1.1102230246251565e-16      1.7320508075688779]   # ■

                                                            <---화면에 있으니 자막에는 없어도 됨)


    이렇게 여러분들은 SVD(singular value decomposition)를 지금부터는 주저 없이 할 수 있습니다. 고등학생과 일반인도 또 수학을 전공하지 않은 사람들도 앞으로는 SVD를 위의 명령어로 바로 구할 수 있습니다.


    Now you can find an SVD of any given matrix by using code above.


    [열린문제 16] 예제 6과 같은 방법으로 다른 교재에서 여러분들이 찾은 행렬의 특잇값 분해를 하라고 물어보는 그런 문제를 만나면, 위의 명령어를 사용해서 구하시면 어려움 없이 바로 찾으실 수 있으실 것입니다.


    [Open problem 16] Find an SVD for a matrix from the other textbook by using the codes above.



    [Review]


    오늘 배운 특잇값 분해를 간단히 요약하자면, 특잇값 분해는 고윳값 대신 특잇값을 이용하여 행렬을 대각화하는 하나의 방법입니다. 이 특잇값 분해가 특히 유용한 이유는 행렬이 정사각행렬이든 아니든 상관없이 모든 크기의 행렬에 대해서 적용 가능하기 때문입니다. 임의 크기(size)의 행렬을 주어도 그것의 SVD를 언제나 구할 수 있다는 의미입니다.


    Singular value decomposition is considered as a method for diagonalizing a given matrix using singular values instead of eigenvalues. It is possible to find an SVD of any sort of rectangular or square matrices.


    이 특잇값 분해는 활용도 측면에서 인공지능에 필요한 선형대수학의 내용 중에서 가장 중요한 내용입니다. 더구나 특잇값 분해에 대해서 위에서 학습한 모든 내용은 정수나 실수 성분의 행렬만이 아니라, 복소수 성분의 행렬에 대하여도 확장할 수 있습니다.


    With a minor modification, singular value decomposition can be done not only for real matrices, but also for any complex matrices.


    즉, 연산이 복소수에서 되게끔 저희가 설정을 실수체가 아니라 복소수 체가 되게끔 코드를 조금 바꿔주기만 하면, 그것도 가능합니다. 왜냐하면 이론이 복소행렬에 다 적용되기 때문입니다. 특잇값 분해를 이용하여 빅데이터 분석에서 주어진 거대한 크기의 행렬을 주요 성질을 보존하면서 인공지능이 쉽게 다룰 수 있는 원하는 크기의 적당한 행렬로 바꾸는 것이 가능하게 하는 것이 이 SVD(singular value decomposition)입니다.


    When we want to have SVD on a matrix with complex entries, we can change the code a little bit, so it can handle computations in complex numbers.


    아까 보셨듯이, SVD를 하시면, 행렬 A가 다음과 같이 직교행렬과 직교행렬 빼면, 이 주대각선 성분의 행렬이 되는데, 크기 순서대로 배열했기 때문에 여기는 크고, 점점 작아져서 마지막 것은 0에 가까워지게 됩니다.


    As you learned earlier, the main diagonal entries of the diagonal matrix Sigma_1 in SVD are arranged in descending order, so the bottom end number is getting closer to zero.


    특히 이것을, 예를 들어서 특잇값(singular value)중 가장 큰 값으로 이 특잇값들을 각각 나누어준다면, 가장 큰 특잇값(singular value)의 값이 1로  변하고, 그러면 그 다음 값들은 1보다 작은 값들이 되고, 점점 더 1보다 작게 돼서, 마지막 것은 0에 가깝게 됩니다.


    If all singular values are divided by the largest singular value, the first main diagonal entry will become 1, and the last main diagonal entry will be very close to zero when the size of the matrix is big enough. The power of the matrix will heavily depend on the several main diagonal entries on the top left.


    예를 들어서, 10,000×10,000 행렬 또는 2백만×백만 행렬같이 큰 크기(size)의 행렬이 된다면, 0아닌 특잇값들이, 예를 들어서, 50만개 있다고 하면, 그 50만개 중에서 가장 큰 특잇값으로 모든 특잇값을 나눠주면, 가장 큰 값이 1이 될 것이고, 그 다음 50만개를 크기순으로 배열하면 아래쪽의 몇 십 만개는 거의 0에 아주 가까운 값이 될 것입니다.


    The same logic can be applied for a very large-size matrix.


    그러면 실제 이 전체 행렬의 행태(behavior)를 이해하는데, 가장 큰 특잇값(singular value)들이 그 분석에 대부분 기여하게 되고, 50만개의 영아닌 특잇값(singular value) 중에서 앞에 있는 몇 천개 또는 몇 백개 또는 몇 십개의 큰 특잇값들이 어떻게 행동(behave) 하느냐가 바로 원래 행렬 A의 행태(behavior)를 설명할 수 있기 때문에, SVD를 한 행렬의 앞의 몇 백개 정도의 특잇값과 이에 대응하는 몇  백개 정도의 중요한 좌특이벡터(left singular vector)와 우특이벡터(right singular vector)들만 이용해서 이 행렬의 앞으로 행태(behavior)를 예측할 수 있게 하는 겁니다.


    To understand the behavior of a given matrix, the largest singular value does contribute a lot. Therefore, the matrix can be down-sized with only a few important left and right-singular vectors corresponding to the large singular values.


    이 과정이 바로 인공지능에서 사용하는 차원축소(Rank Reduction, dimension reduction)이라는 개념입니다.


    This idea is called a “rank reduction” or a “dimension reduction.”


    인공지능에서 엄청난 크기의 거대 빅데이터 행렬을 인공지능이 쉽게 다룰 수 있게, 우리가 원하는 적당한 크기의 행렬로 바꾸어서, 컴퓨터가 바로바로 계산해서 답을 줄 수 있도록 하는 그 모든 이면의 얘기(behind story)가 바로 이 SVD(singular value decomposition)를 활용해서 이루어지고 있다는 것을 이해하시면, 왜 SVD가 인공지능의 핵심(Core, key) 아이디어인지를 이해하실 수 있으실 겁니다.


    행렬 분해에 대한 개념, 구체적인 유도 과정 및 응용 사례는 고등학교, 일반인들의 수준을 넘을 수 있습니다. 그래서 그 동안 자세히 가르치지를 않았는데, 지금은 꼭 필요한 내용이라서 아주 쉽게 설명해드렸고, 실제 SVD를 할 수 있는, QR분해를 할 수 있는 방법 및 코드를 알려드렸습니다.


    In AI, SVDs are really needed to easily handle big data matrices in a short time by the rank-reduction technique.


    이제 고윳값과 행렬의 대각화 및 SVD의 배경이 되는 지식에 대해서 더 궁금하신 분들은 선형대수학 교재나 웹 교재 또는 실습실에 있는 고윳값 부분, 행렬의 대각화 부분을 학습하시면 됩니다.


    More details on matrix diagonalizations can be found in the following link: http://matrix.skku.ac.kr/LA-K/


    더 자세한 내용은 제가 대학생용으로 책으로 쓴 <인공지능을 위한 기초수학> 교재를 무료로 다운(down) 받아서 참고하시면 됩니다. 모두 [온드림 BigBook을 통하여 무료로 제공한] 전자책이니까, 바로 다운(down)받아서 사용하시면 됩니다.


    그리고 오늘 설명한 부분을 조금 보충하면 LU분해 도구를 이 주소에 따로 만들어두었으며, QR분해 도구 및 최소제곱해에 대한 도구도 따로 실습실을 만들어 두었으니까 이용하시면 됩니다.


     실제 데이터로부터 Ax=b라는 선형연립방정식을 정의하고, 그 문제의 해를 구해보면, 많은 경우 해가 존재하지 않는 것을 확인할 수 있습니다.


    해가 존재하지 않지만, 그 때 우리가 Ax와 b의 차이가 가장 작게 되는 <그 문제의 최소제곱해>를 QR분해를 이용해서 구하는 과정을 자세하게 설명해주었고, 그리고 수업시간에 설명했듯이, 문제의 조건이 우리가 사용할 알고리즘을 사용할 수 있는 조건에 부합하는지를 먼저 확인했고, 그 다음 코드를 이용하여 행렬 Q와 행렬 R을 쉽게 구해서 실제 Q, R을 곱하면 A 하고 같아지는 것을 확인했습니다. 그리고 또 다른 방법으로 solve 명령어를 이용하여, 구해보면 다음과 같이 나오는 것을 확인할 수 있습니다.


    A system of linear equations that has no solution, can also be handled by minimizing the distance between Ax and b. This process can also be done the using QR decomposition. If you actually multiply Q and R, you can see that the difference between A and QR is almost zero. 


    이렇게 구한 답을 이전에 우리가 다른 방식으로 배웠던 방법, 예를 들어 정규방정식(normal equation)을 풀어서 구한 답하고 비교해보면, 같은 답이라는 것을 확인할 수 있습니다.


    그리고 A의 열(column)들이 일차독립이 아닐 경우에는 어떻게 전처리를 하는지, 그 테크닉은 아래 주소에서 확인하시면 됩니다.


    If the columns of A are not linearly independent, we can only find linear independent columns before we start the algorithm. More details can be found in the following link: http://matrix.skku.ac.kr/LA-K/Ch-8/


    또, SVD를 실습하실 수 있도록 특잇값 분해는 별도 웹 주소의 실습실을 만들어놨습니다. 이 주소에는 본인이 50분 정도 자세히 설명한 내용과 특잇값 분해에 대한 기초이론들을 다시 정리해주었습니다.


    실습실에서는 실수체  RDF 대신, 복소수체 CDF로 정하면 complex까지도 포함할 수 있습니다. 그래서 필요시 항상 SVD를 구해서 U와 와 V를 구해서 A하고 와의 차이(error)가 얼마나 되는지 비교해 보시면, 거의 차이(error)가 없다는 걸 확인할 수 있습니다. 즉 우리가 코드로 구한 값이 거의 맞다는 것도 확인할 수가 있었습니다.


    In the code for finding an SVD, RDF works for real matrices, and CDF works for complex matrices.


    즉 SVD 분해에 대한 도구는 위에 언급한 웹 주소에 대부분 제공했으니까 확인해 보시고,  SVD 이론을 더 알고 싶으신 분은 이 자료와 동영상에 더 자세한 설명이 있으니 확인하시면 됩니다.



    [Review] 


    이제 6주차에 배운 내용을 복습하도록 하겠습니다.


    이번 6주차에는 인공지능에 가장 필요하고 중요한 이론적 배경인 SVD(singular value decomposition)와 최소제곱해(least square solution)를 가장 쉽게 구하는 QR분해에 대해서 배웠습니다.


    In today’s class, we have learned some important matrix decompositions, including LU-decomposition, QR-decomposition, and SVD.


    이번 6주차 수업에서는 우리가 인공지능에 가장 많이 사용되는 행렬분해 기법인 LU분해와 QR분해, 그 다음에 SVD를 배웠습니다.



    이 QR 분해는 최소제곱해(least square solution)와 직접적인 관계가 있었고. Gram-Schmidt 정규직교화법과 관계가 있었고, 특잇값 분해는 행렬의 대각화의 일반화된 모양이라고 이해하시면 됩니다. LU분해는 행렬 A가 LU로 분해되면 Ax=b라는 연립방정식을 푸는 문제가 엄청 쉬운 문제로 바뀐다는 것을 이해하셨고. 행렬 A를 QR로 분해할 수 있으면, Ax=b라가 문제가 해가 존재하지 않을 때, 그 최소제곱해(least square solution)을 구할 때에도 바로 Q와 R만 대입해서 최소제곱해를 구할 수 있다는 것을 배웠으며 이 특잇값 분해에서는 행렬 A를 로 고쳐주면, 이 , 이것을 봐서, 이것의 주대각선 성분인 특잇값(singular value)들의 개수를 우리가 원하는 만큼 줄여서 이 행렬 크기(size)를 적당히 적은 문제로 바꿔주고, 이 적은 문제를 이용해서 문제를 해결해서 얻은 답이 우리가 큰 크기(size)의 문제에 대한 답과 큰 차이가 없는 의사 결정을 할 수 있으면서도 계산 시간은 거의 1,000만 배, 100만 배만큼 줄일 수 있다는 그런 내용들을 학습 했습니다.


    QR-decomposition is related to the least squares problem, and the SVD is a generalization of a matrix diagonalization.


    LU decomposition helps us to easily solve the equation Ax=b when a matrix A are decomposed into LU. QR-decomposition can be used even when there is no solution for Ax=b. The SVD can be obtained for all rectangular matrices A.


    지금까지 매우 중요하고도 핵심적인 내용들을 학습해 오셨는데, 혹시 그동안 배운 내용들을 요약한 것을 보고 싶으면 제가 가르치는 학생들이 그 동안 배운 내용들을 요약한 아래 주소의 보고서를 참고하시면 됩니다.


    수고하셨습니다. 다음 시간부터는 <미적분학>을 다루며, 극한과 도함수로 시작해서 인공지능에 필요한 경사하강법(gradient descent method)을 소개하도록 하겠습니다.


    From the next class onward, we will go through Calculus and the gradient descent method. Thank you.






    Week 7. Introductory Mathematics for Artificial Intelligence

     

                인공지능을 위한 기초수학 입문


    III.  인공지능과 최적해 (미분) (7-9주차)     75

    7. 극한과 도함수 75

    7.1 함수의 극한 (limit)

    7.2 도함수(derivative)와 미분(differentiation)


    [7-1 pre]


    안녕하십니까. Math4AI 7주차 강의는  ‘극한과 도함수’입니다.


    Hi!!  ‘Limits and Derivatives’ is the to.\PIC for this week.


    이번 주에는 함수의 극한과 도함수에 대해서 학습합니다.


    We will study <the limit and derivative of a function>.


    인공지능은 데이터를 바탕으로 모델을 만들고 가장 적합한 파라미터와 해를 찾기 위하여 최적화 과정을 거치게 됩니다. 이는 극한과 도함수의 개념을 필요로 합니다.


    Artificial intelligence creates a mathematical model based on Data and goes through an optimization process to find the best possible parameters and solutions. This requires the concept of limits and derivatives.


    이번 시간에는 함수의 극한과 미분, 그리고 도함수의 개념에 관하여 학습합니다.


    Today, we will study the limits and derivatives of functions.


    1절에서는 함수의 극한, 2절에서는 도함수와 미분에 설명합니다.


    <Section 1> covers the limits of functions, and <Section 2> covers the differentiation and the derivative.


    함수의 극한은 고등학교 때, 이미 배운 내용입니다.


    The limit of a function is what we already learned in high school.


    함수의 극한을 이용하면, 함수가 어떤 점에서 수렴하거나 발산하는지 확인할 수 있고, 도함수를 극한을 이용하여 정의하며, 그 도함수를 구하는 미분을 복습합니다.


    Using the limit of a function, we can check whether a given function converges or diverges at a certain point, and we can define the derivative of a function. We will review the differentiation.


    그 후 도함수를 이용하여 한 점에서 접선의 방정식을 구할 수 있습니다.


    With the derivative of a function at a point, we can find the equation of the tangent line at that point.


    [7주차 1강] 반갑습니다. K-MOOC [인공지능을 위한 기초수학 입문], 7주차. <인공지능과 최적해>. 여러분들이 익숙한 미분법, 도함수에 관련된 내용입니다.


    [Week 7, Lec 1] In this seventh week, we will cover <Calculus>, including the concepts of differentiation and derivatives on single variable functions.


    미적분학의 개념에 대해서는 제가 말로 풀어 쓴 ‘스토리텔링 Calculus’를 보시면, 미적분학의 역사와 관련 스토리를 확인하실 수 있을 것입니다.

     http://matrix.skku.ac.kr/Calculus-Story/index.htm  


    For <History and Concepts of Calculus>, please refer to the link below:

    http://matrix.skku.ac.kr/Calculus-Story/index.htm 



    이번 주 첫 시간에는, <함수의 극한과 도함수>에 대해서 학습하겠습니다. 근사해를 구하는 것과 최적해(optimal solution)를 구하는 문제는 관련성이 큽니다. 가능한 최선의 해(best possible solution)를 구하는 겁니다. 최적해를 구하는 문제는 도함수의 일반화된 개념과 연산들을 포함합니다. 그래서 극한과 도함수 개념의 복습으로 시작합니다. 이 내용은 아래 주소에서 실습하실 수 있습니다.

               http://matrix.skku.ac.kr/math4ai-intro/W7/


    In the first session, we learn about <Limit and Derivative of the Function>. It is related to the problem of finding an approximation or optimal solution for any given problem. The problem of finding an optimal solution includes definition and computation on the derivative of functions. So let's start with the definition of limit and derivative. You can practice them at the following link.

      http://matrix.skku.ac.kr/math4ai-intro/W7/


    Section 1.  Limit                1절 함수의 극한


    먼저 다음 질문을 생각해보죠.

    가 에 한없이 가까이 갈 때(1에 한없이 접근해갈 때), 은 어떤 값을 가지게 될까요? 이 경우 는 일 때는 분모가 이 되므로 함숫값이 정의되지 않습니다. 그러나 이면,  만 남으니까 한 점을 제외하고는 이라는 그래프 모양을 갖게 됩니다. 여기 구멍이 있습니다. 따라서 의 그래프는 오른쪽 그림과 같이 의 값이 1이 아니면서 1에 한없이 가까워질 때, 의 값이 2에 한없이 가까워지는 것을 알 수 있습니다.


    Let’s consider the following question.

    If x approaches 1, then what will happen to the value of ?

    In this case,  has a zero denominator when . Hence this function cannot be defined at . In case of , it can be rewritten as , so you can see that it is except for one point, x=1. Therefore, The graph of can be seen as x approaches 1 but not equal to 1, the value of approaches 2.


    이때, 코딩을 이용하여 의 극한을 직접 확인할 수 있는데, 식으로 계산하거나 그래프 그리기 활동을 코딩으로 쉽고 간단하게 예측하고 그 값을 구할 수 있으며 이 방법을 이후 다양한 함수로 확장하는 것도 가능합니다.


    Now we can compute the limit of by using a simple Sage code, which is much simpler than drawing a graph by hand. This code can be used for various functions to be drawn.


    ---- http://matrix.skku.ac.kr/KOFAC/ --------------------

    f(x) = (x^2 - 1)/(x - 1)  # 함수 정의

    show(plot(f(x), (x, -0.5, 1.5), ymin = -1, ymax = 2.5, detect_poles = 'show', exclude=[1]))

    # 함수 그래프,  x=1에서 극한값 그래프를 확인하기 위해 범위 주어 선택

    n = 10

    for i in range(1, n):   # 1보다 크면서 1에 가까이 있는 x값에 대해서 f(x) 계산

        s = (1 + 1/(10^i)).n()

        print("x =", s, "f(x) =", f(s).n())

    print()  # 한 줄 비우기

    for i in range(1, n):   # 1보다 작으면서 1에 가까이 있는 x값에 대해서 f(x) 계산

        s = (1 - 1/(10^(n - i))).n()

        print("x =", s, "f(x) =", f(s).n())

    -----------------------------------------------------------------


    위의 명령어를 이용하여 1 근처에서의 함숫값 들을 계산하여 그래프를 그린 것입니다. ‘n=10’은 10개의 점에서 그 구간을 1/10로 나눠서 10개의 점에서의 를 구하라는 의미입니다.


    We can find the values of at some points near 1, and draw a graph. In the above codes, ‘n=10’ means dividing the interval by 10.


    직접 구해보면, 먼저 일 때 는 2.1이고, 1.01일 때는 2.01이고, 가 1.000001일 때는 2.0000001입니다. 따라서 x가 1에 한없이 가까워지면, 는 거의 2에 수렴하는 것을 확인할 수 있습니다.


    The exact values are given as follows.

    =2.1 for

    =2.01 for x =1.01

    =2.0000001 for x =1.000001

    Therefore, when x approaches 1, converges to 2.

    ...

    x = 1.00000100000000 f(x) = 2.00000100008890

    ..

    x = 1.00000000100000 f(x) = 2.00000000000000


    x = 0.999999999000000 f(x) = 2.00000000000000

    ...

    x = 0.999990000000000 f(x) = 1.99998999999917

    ...


    다음과 같이 x가 양방향에서, 는 모두 2로 수렴하는 것을 볼 수 있습니다.


    In both (left and right) directions of x, converges to 2.


    이와 같이 파이썬 기반의 코딩언어를 이용하여 직관적으로 가 1에 수렴할 때 는 에 수렴함을 확인하고 예측할 수 있었습니다. 이것을 간단한 명령어


     limit((x^2 - 1)/(x - 1), x = 1)  # x =1로 수렴할 때, f(x)의  limit f(x)를 구해라.


    By using our Python-based Sage codes, we can intuitively predict that converges to 2 when x approaches 1.


    를 이용하여 계산하면 바로 2가 나옵니다. 간단하게 구할 수 있습니다.


    따라서 함수 는 의 값이 1이 아니면서 1에 한없이 가까이 갈 때, 2에 가까워지는 것을 확인할 수 있습니다. 이를


    일 때 또는

    와 같이 씁니다.


    You can see that when the value of x approaches 1, the function converges to 2. We denote this by as or simply  .



    이처럼, 가 로 수렴할 때 가 에 수렴하면, “x가 에 가까이 갈 때, 는 에 수렴한다”고 하고 로 표기합니다. 이때 를 의 극한(limit)라고 합니다. 이 극한이 존재하지 않으면, 즉 수렴하지 않으면 발산한다고 합니다.


    If converges to as approaches , then we say this "As x approaches , converges to " and denoted by . This is called the limit of . We say diverges if the limit does not exist.


    이제 앞서 계산한 함수 를 다양하게 바꿔보고 그 극한값을 구해볼 수 있습니다. 아래의 예에서, 가 근처에서 발산함을 그래프를 그려서 확인할 수 있습니다.


    You can change the function f(x) and find its limit at any point. In the example below, the graph shows that  diverges at .



    ------------------------------------------

    f(x) = (1)/(x^2)  # define a function

    show(plot(f(x), (x, -5, 5), ymin = -2, ymax = 2, detect_poles = 'show'))

    ------------------------------------------------


    위의 명령어를 쓰면, 의 그래프가 다음과 같이 그려지는 걸 확인할 수 있습니다. 또 x가 0에 수렴할 때 가 발산하는 것을 수치적으로도 확인할 수가 있습니다.


    Using the code above, we draw the graph of as follows. It is also possible to verify numerically that diverges when x=0.


    [참고]  위에 제시된 극한의 정의는 고등학교에서 배우는 직관적인 내용입니다. 그래프를 그리면 어디로 수렴하는지 바로 확인할 수 있습니다. 그렇지만 극한의 엄밀한 정의는 고등학교 수준을 넘어섭니다. 실제로 극한의 개념을 완전히 이해하는데 전문수학자들도 150년이라는 시간이 걸린 내용입니다. 이 내용이 (Epsilon-Delta) 정의를 이용해서 증명이 되었습니다. 그 자세한 내용은 미적분학 chapter 2에서 확인하실 수가 있습니다. (참고)  http://matrix.skku.ac.kr/Cal-Book1/Ch2/


    [c.f.] The definition of limit described above is an intuitive one. If you draw a graph, you can see where it converges. A precise definition of limit requires the Epsilon-Delta argument. For more information, please refer to Chapter 2 in Calculus. http://matrix.skku.ac.kr/Cal-Book1/Ch2/


    다음은 앞에서 설명한 함수 의 우극한(right limit)과 좌극한(left limit)을 직접 계산해 보겠습니다. 


    Let's compute the right-hand limit and left-hand limit of the function described earlier.


    [예제 1] 다음 극한(limit)을 구해봅시다.


    [Example 1] Find a limit of the following functions.


    1)  (limt) ,  Answer 1)  = .

    2)    (left-hand limit),    Answer 2)  = -4.

    3)  (right-hand limit)    Answer 3) = 4.


     Check them with the following Sage codes.


    ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    var ("x")

    f(x)=(x^2-4)/abs(x-2)

    show (f(x))

    show ( plot(f(x), (x, -2, 4), ymin = -5, ymax = 5, exclude=[2]))

    print(limit(f(x), x = 2, dir = '-'))  # left-hand limit at x = 2

    print(limit(f(x), x = 2, dir = '+'))  # right-hand limit at x = 2

    -------------------------------------------------------------------

    Output : 

      -4      # left-hand limit at x = 2

       4       # right-hand limit at x = 2   


    음(minus) 방향으로 2로 수렴할 때는 함숫값(좌극한) 이 –4 값으로 수렴하고, 양(positive) 방향에서 2로 접근하면, 함숫값(우극한)이 4 로 수렴하게 됩니다. 절댓값 함수니까 그렇습니다.


    When x approaches 2 from the left, the value of function converges to -4, and if x approaches 2 from the right, the value of function converges to 4. This is because it was an absolute value function.


    우극한 와 좌극한 가 모두 존재하고, 그 값이 같으면 에서의 극한값이 존재한다고 합니다. 앞의 예 경우, 각각 좌극한과 우극한이 존재하지만, 하나는 -4, 하나는 +4로, 다르기 때문에 그 점에서 limit가 존재하는 것은 아닙니다.


    If both the right-hand limit and the left-hand limit exist, and if the values are the same, we say that there exists a limit of at .


    In the previous example, the limit does not exist because the two values are different.


    [예제 2] (1) 와 (2) 를 구하는 문제입니다.

            [Example 2] Find, (1) and (2) .


            풀이.  를 구하기 위해, 좌극한과 우극한을 구하여 비교한다.

                  이면 이므로, ,

                  이면 이므로,

                  이므로 의 그래프는 그림과 같다.

                  따라서 로 양의 무한대로 발산한다.

    Sol) Find and compare, the left and right-hand limits to find .

    If , then . So, .

    If , then . So, .

    The graph of is given as follows.


    It is easy to understand when we draw the graph.


    ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    f(x) = sin(x)/abs(cos(x))

    show(plot(f(x), (x, -2, 4), ymin = -5, ymax = 5, detect_poles = 'show'))

    print(limit(f(x), x = pi/2))  # limit at x = pi/2


    g(x) = x^2/sin(x)

    show (g(x))

    show(plot(g(x), (x, -2, 4), ymin = -5, ymax = 5, detect_poles = 'show'))

    print(limit(g(x), x = pi/4))  # limit at x = pi/4

    ---------------------------------------------------------------------


    (1)  +Infinity  # No limit at x = pi/2 . diverges  .


    (2) 1/16*sqrt(2)*pi^2    # limit exists, converges.     ■




    그래프를 보면 쉽게 이해가 됩니다.


    첫 번째 그래프는 pi/2에서 함수값이 무한대이므로 pi/2에서 함수의 극한이 존재하지 않아서, 발산하는 경우입니다. 두 번째 그래프는 다음과 같으며, 주어진 값에서 극한이 으로 존재해서 수렴하는 예입니다.


    (1) The graph has an infinite value at π/2, so there is no limit.

    (2) The graph is an example of convergence because the limit exists as .



    * 함수의 극한에서는 일 때 값의 존재여부가 중요하지 않다. 즉 가 존재하지 않아도 극한 은 정의될 수 있다. 만일 가 정의되고 이면 는 에서 연속(continuous)이라고 합니다.

    즉 한 점에서 극한(limit)이 존재하고 그 점에서 값이 존재해서, 그 둘이 같으면 연속이라고 합니다.


    Even for the case that does not exist,  may be defined. If is defined and , then we say that is continuous at .


    이상 우리가 배운 내용을 복습하겠습니다.

    7주차 1차시 <극한과 도함수>에서 극한(limit)을 학습하고 의 그래프를 그리고, x=1 근처에서 함수들이 어떻게 수렴하는지를 확인했습니다. 직관적으로 x가 1에 가까이 갈 때, 는 2에 다가가는 것을 확인하였고, 그것을 실제로 명령어 ‘limit’ 를 주어서 값이 같음을 바로 확인할 수 있었습니다. 또 함수 에서 x가 0에 수렴하면, 함숫값은 무한대로 발산한다는 것을 확인했습니다.

    또 x가 2에서, 이 함수의 좌극한과 우극한을 구해서, 좌극한과 우극한이 같으므로 극한(limit)가 존재하고, 또 도 존재해서 이 둘이 같으므로 이 함수는 x=2에서 연속임을 보였고, 그러나 x가 –1에서는 발산한다는 것도 (그래프에서) 확인할 수가 있었습니다.


    [Review] 


    In this first session, we have learned how to draw a graph of , how to find the left and right-hand limits, and how to check whether the given function converges or diverges near a point x. Intuitively, when x approaches 1, converges to 2. If the values of the left and right-hand limits are same at a point, the function has the 'limit' at that point. It was also shown that the function diverges at x=0.

    Limit of exists, because the left and right-hand limits are same at x=2. The function is continuous at x=2, because also exists and the two values are the same. But it diverges at x=-1.



    다음 경우는 좌극한과 우극한이 다르기 때문에, 극한(limit)이 존재하지 않으며, 따라서 연속도 아닙니다. 그리고 함수 의 그래프를 그려봤습니다. 그래서 이 함수는 에서는 (무한대로) 발산한다는 것을 확인했고, 구간을 축소하여 x=1 에서는 어디로 수렴하는지 확인했습니다.


    In the following cases, the limit does not exist, and moreover, the function is not even continuous, because the left and right-hand limits are different. We could see the function     diverges at x=, and it converges at x=1 by adjusting the interval in our plotting with the code.


    이제, 다음 2차시에는 도함수(derivative)와 미분(differentiation)에 대해서 학습하겠습니다.  수고하셨습니다.

    In the second session, we will study more on the derivative and differentiation. Thank you.


    [7-2 pre]


    안녕하십니까. 인공지능을 위한 기초수학 입문. 7주 2차시에서는 앞에서 배운 극한과 도함수 개념을 이용하여 도함수와 고차도함수를 구하고, 접선을 구하여, 그래프를 그리는 실습을 합니다.


    Hi!! Welcome to the <Introductory Math for AI> class. In this second lecture of Week 7, we will practice how to draw a graph of the tangent line at a certain point after we find the derivative and the higher order derivatives of a given function.


    [7주차 2강]  반갑습니다. 7주차 2차시 <도함수와 미분>입니다.

                http://matrix.skku.ac.kr/math4ai-intro/W7/


    [Week 7, Lec 2] Lecture 7, the second session, <Derivative and Differentiation>. http://matrix.skku.ac.kr/math4ai-intro/W7/


    지난 시간에는 극한(limit)과 연속에 대해서 학습을 했습니다.

    이번 시간에는 도함수(derivative)와 미분(differentiation)에 대해서 학습합니다.

     

    In the first session, we covered the limit and continuity of functions. Now we will talk about <Derivative and Differentiation>.


    먼저 아래 질문으로 시작합니다.


     한 점 에서 포물선 에 대한 접선의 방정식을 구하라.


    Let's start with the following question.


    Q) Find an equation of the tangent line to at the point .


    이 때 함수 위의 점 에서 접선의 기울기가 이라고 알고 있다면 다음 공식에 의해 <접선의 방정식>을 쉽게 구할 수 있습니다. 점이 주어지고 기울기가 주어지면, 자연스럽게 그 점에서의 접선을 다음과 같은 방정식으로 표현할 수 습니다.


    고등학교 내용인 접선의 방정식에서 배운 내용입니다.


    If the slope of the tangent line is given by at the point of the function , then the tangent line can be obtained by the following formula.

                                 .


    우선, 포물선 위의 점 근처에서 한 점 를 선택하여 할선(secant) 의 기울기를 한번 계산해봅시다. 할선을 secant라고 하는데 이는 기울기를 의미하며, (1, 1) 점에서는 가 기울기가 됩니다. 할선(secant)은 자른다는 의미인데, 그래프의 이 두 점 사이를 자른 이것이 기울기입니다.


    Choose a point near the point on the parabola . Then compute the slope of the secant line . At (1,1) is the slope.


    그림으로 보시면, 점 를 점 로 이동함에 따라서 할선 는 접선에 그 한 점에서 접선에 점점 가까워지는 지는 것을 이해할 수 있습니다. 따라서 접선의 기울기는 할선 의 기울기의 일 때 이 값의 극한으로 정의하는 것이 타당합니다. 이렇게 해서 도함수를 정의합니다.


    In the figure, as the point moves to point , it is easy to see that the secant line is getting closer to the tangent line. Therefore, the slope of the tangent line is defined as the limit of the slope of the secant line as . Now we define the derivative with this concept.


    즉,  입니다. 따라서 포물선 위의 점 에서의 접선의 방정식은 입니다.  즉, 라는 접선이 그려지는 것입니다.


    So, =  .


    Hence the equation of the tangent line to at is , which means .


    함수 의 정의역 내에 속하는 점 에 대하여, 다음과 같은 극한값이 존재하면 함수 는 에서 미분가능(differentiable)하다고 합니다. 그리고 이 극한값을 에서의 함수 의 미분계수(differential coefficient)라 하며 라고 씁니다.


    For a point in the domain of the function , if the following limit exists, then the function is said to be differentiable at . This limit value is called the derivative of function at and is denoted by .



    앞서 언급한 접선을 생각하면 미분계수는 접선의 기울기를 나타냅니다.


    Considering the tangent line mentioned above, the derivative indicates the slope of the tangent line.


    따라서 라는 점에서 의 접선 방정식은 다음과 같이 m 대신에 미분계수를 이용해서 사용할 수 있습니다. 점 (, )에서, 미분계수  가 들어가면 이 되는 것입니다.


    Therefore, an equation of the tangent line to at  can be expressed using the derivative instead of m, as follows:

                          .


    는 에서 미분가능이면 에서 바로 연속입니다. 미분가능이면 연속 (그리고 미가연적- 미분가능이면 연속, 연속이면 적분가능). 그러나 역은 성립하지 않습니다. 즉 가 에서 연속이라도 반드시 미분가능인 것은 아닙니다. 그림을 보시면, 이 점에서 연속이 아닙니다. 그리고 연속이 아니니까 당연히 이 점에서는 미분이 불가능합니다.


    If is differentiable at , then must be continuous at . But the converse is not true. In the .\PICture, is not continuous at this point, which means is not differentiable at the same point.


     이 그림은 함수가 연속이면서 미분가능한 구간, 연속이지만 미분가능하지 않은 점, 기울기가 무한대라 미분가능하지 않은 점 등을 보여줍니다. 그래서 이 하나의 그래프는 <연속이고 미분가능인 경우, 도함수가 무한대인 경우, 연속이지만 미분 불가능한 경우, 연속이 아니므로 미분이 불가능한 경우>를 모두 보여주는 좋은 예입니다.


    This figure shows an interval that the function is continuous and differentiable, a point that the function is continuous but not differentiable at the point, and a point that the function is not differentiable and the slope is infinite.


    가 어떤 구간의 각 점 에서 미분가능일 때, 는 이 구간에서 미분가능이라고 합니다. 그리고 이 경우 각 점 에 그 점에서의 미분계수를 대응시킴으로써 정해지는 함수를 의 도함수(derivative)라고 합니다. 도함수는 기호로는 다양하게 표시되죠. 도함수를 표시합니다.


    When is differentiable at every point x in an interval, then is called differentiable in that interval. In this case, the derivative at that point is called the derivative of f(x) at x. Denoted by  .


    함수 의 도함수를 구하는 것을 를 미분(differentiate)한다고 얘기하고, 함수 의 도함수는 = 로 구합니다. 따라서 어떤 함수를 미분하여 얻은 그 함수가 도함수이고, 거기에 변수의 값을 대입하면 그 점에서의 미분계수가 나오는 것입니다.


    Finding the derivative of a function is called the "differentiation of ." The derivative of a function can be found as = . Therefore, the function obtained by differentiating a function f(x) is called the derivative f'(x).


      함수 의 도함수 가 다시 미분가능이면 그 도함수의 또 도함수를 구할 수 있죠. 그것을 우리는 y의 2계 도함수(2nd derivative)라고 하고, 등으로 표시합니다. 2계 도함수가 또 다시 미분가능이면 3계 도함수를 생각할 수 있고, 이런 식으로 4계, 5계, 계속 번 미분하면 계 도함수가 정의되므로, 계 도함수(-th derivative)를 다음과 같이 로 표현합니다. 


    If the function of is also differentiable, then the derivative has its derivative (f'(x))' again. This derivative is called the second derivative of y. Denoted by . If the second derivative  is also differentiable, then the third derivative can be found and so on. We denote for -th derivative of y=f(x).


      n계 도함수가 존재하는 경우에 는 번 미분가능하다고 얘기합니다. 미분계수의 응용인 속도와 가속도는 각각 거리를 나타내는 함수의 1계 도함수와 2계 도함수의 예입니다. 속도와 가속도. 연속인 함수가 있으면 1계 도함수는 그 함수의 속도를 얘기하고, 2계 도함수는 그 점에서의 가속도를 우리가 생각할 수 있게 합니다.


     If the -th derivative exists, is said to be times differentiable. Speed and acceleration, as applications of the derivative, are the first derivative and second derivatives of the distance function, respectively.


    미분에 관한 기본 성질들, 정리가 다음과 같습니다. (화면의 수식을 보세요)


     Fundamental properties of a derivative.


     ⓵

     ⓶

     ⓷

     ⓸ , 단


     아주 중요한 식으로 고등학교 때 배우는 내용입니다.


    These properties are already covered in high school math.


    [참고]  그 밖에 여러 가지 미분 공식에 관한 자세한 내용은 아래에 있습니다. 

              http://matrix.skku.ac.kr/Cal-Book1/Ch3/


    More information on derivatives can be found in

     http://matrix.skku.ac.kr/Cal-Book1/Ch3/ 



    이제 실제 y 가 다음과 같이 주어진 미분가능한 함수일 때, 도함수와 2계 도함수, 3계 도함수를 구해보겠습니다.


    For a given differentiable function y, let's find the first derivative, the second derivative, and the third derivative.


    [예제 3] 의 도함수(derivative)와 2계 도함수 , 3계 도함수 를 구해보겠습니다. 우리는 손과 code를 이용해서 구합니다. (화면의 코드를 보세요)


    [Example 3]  Let . Find, , , and .


    Sol.  , ,


    ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    f(x) = x^2 + 4*e^x

    print(diff(f(x), x))   # dy/dx

    print(diff(f(x), x, 2))  # second derivative

    print(diff(f(x), x, 3))  # third derivative

    ---------------------------------------------------------------------

    2*x + 4*e^x      # dy/dx, 도함수  

    4*e^x + 2        # 2계도함수

    4*e^x             # 3계도함수                                      ■

                                                             <---화면에 있으니 자막에는 없어도 됨)


    실제 명령어로 구해보니까, 첫 번째 도함수가 일치합니다. 2계 도함수도 2+4*e^x 로 일치합니다. 3계 도함수도 4*e^x 로 정확히 일치합니다.


    If you use the code, you can differentiate various differentiable functions.


    이렇게 손으로 구하던 도함수를, 코드를 이용해서 주어진 함수가 무엇이건 간에, 여러분들이 미분가능한 함수를 입력하면, 또 3차 미분가능한 함수를 이곳에 대입하면 어떤 함수에 대해서도 답을 정확히 얻을 수 있습니다.

     


    [예제 4] 의 도함수(derivative)와 2계 도함수 를 구해봅시다.


    Let's do more exercise problem.

    [Exercise 4] For a function , find , and .


    즉, ‘다음과 같은 함수의 도함수와 2계 도함수를 구하라’로 할 때, 주어진 함수가 점점 복잡해지고, 합성함수도 들어갈 수 있습니다. 그러면 미분법과 곱의 미분법의 규칙들을 이용해서 1계 도함수 구하고 2계 도함수를 이렇게 구해야 됩니다. 이것이 3계도함수, 4계도함수로 가면 점점 복잡해지고, 함수가 복잡해지면 도함수도 점점 복잡해집니다. 그런 경우 손으로 구하기보다는, 주어진 함수가 미분가능한 함수인지 아닌지만 확인하고, 미분가능하면, 앞에서 했듯이, 1계 도함수, 2계 도함수 명령어를 활용해서 이 함수를 입력하고 명령어를 클릭합니다. (화면의 코드를 보세요)


    When we were asked 'To find the first derivative and the second derivative of the following functions', the given functions may be complicated and possibly contain composite functions. Using differentiation rules, the 1st and 2nd order derivatives can be obtained. But if you have to find the third and fourth derivatives, it becomes more and more complicated, and if the functions are very complicated, finding derivatives becomes more and more difficult. In these case, rather than obtaining it by hand, it is better to use codes after we check that the given function is differentiable. The following codes and commands will be useful.


    -------------------------------------------------------------------

    f(x) = sin(x^3) + 4*(e^(x^2))

    print(diff(f(x), x))   # dy/dx

    print(diff(f(x), x, 2))  # 2계도함수

    ------------------------------------------------------------------

    3*x^2*cos(x^3) + 8*x*e^(x^2)                 #  f'(x)

    6*x*cos(x^3) -9*x^4*sin(x^3) + 8*e^(x^2) + 16*x^2*e^(x^2) # f''(x)  ■

                                           


    클릭만하면 1계도함수와 2계도함수는 바로 나옵니다. 이 방법이 우리가 손으로 구한 것보다 전혀 못하지가 않습니다. 정확하게 일치하는 답을 구하는 걸 알 수 있습니다. 그리고 매우 빠르게 구해주기도 합니다. 이런 식으로 단순계산은 코드를 이용해서 얼마든지 할 수 있습니다.


    The first derivative and second derivative of functions come out immediately using codes. There is no difference between calculating by hand and codes. Using some codes, simple calculations can be done easily whenever we need to do.



        [예제 5] 곡선 위의 점 에서의 접선의 방정식을 구해봅시다.


        [Exercise 5] Find a tangent line of the curve  at .


        도함수 구하고 접선의 방정식에 대한 공식에 바로 집어 넣어주면 됩니다. 접선을 구하는 문제는 앞으로도 많이 사용해야 되므로 코드를 하나 만들어두었습니다. (화면의 코드를 보세요)


        You can obtain a derivative and put it into the equation of the tangent line. The following codes will help you to get it.



    ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    f(x) = x^2 + sqrt(x)

    x0 = 4

    df(x) = diff(f(x), x)    # derivative

    print("접선의 방정식 :  y =", f(x0) + df(x0)*(x - x0))

    p1 = plot(f(x), (x, 0, 10))   # plot f(x)

    p2 = plot(f(x0) + df(x0)*(x - x0), (x, 0, 10), color = 'red')   # 접선 그리기

    p1 + p2   # plot f(x) and tangent line on a common screen

    -------------------------------------------------------------

    tangent line:  y = 33/4*x – 15  # y``=f``(a`)``+`f`` prime `(a`)``(x-a`)`=`-15`+` {33} over {4} x ■

    The tangent equation at (4,18) is .


     그러면 에서의 접선의 방정식 , 즉 명령어에 의하여 f(x0) + f’(x0)*(x – x0)) 가 정확히 나옵니다. 원래 방정식과 접선의 방정식이 같이 그려지면서 내용을 쉽게 확인할 수 있습니다.



    [열린문제 1] 여러분들은 여기서 배운 것을 이용해서 다른 교재에서 찾은 n계 미분가능한 함수를 가지고 3계 도함수(3rd  derivative)를 구해보는 연습문제를 어렵지 않게 해보실 수 있습니다.


    [Open problem 1] Find the third derivative for the differentiable function from the textbook.


    [참고자료]  7주차에서 배운 내용들과 관련해서 여러 참고자료를 준비했습니다. 관련 코드들과 해를 공유해 놨습니다. 명령어를 바꿔서 더 복잡한 함수를 풀어보고 싶으시면, 아래 주소에서 미적분, 명령어 모음을 확인하실 수 있고 미적분에 대한 강의,, 또 일변수 함수에 대한 미적분, 또 다변수함수에 대한 미적분에 대해 학생들이 학습한 내용들을 볼 수 있습니다.

      그러면, 명령어 모음 http://matrix.skku.ac.kr/Lab-Book/Sage-Lab-Manual-1.htm 를 소개하겠습니다. 함수의 그래프, 함수의 그래프를 그리는 명령어부터 시작해서 그래프의 식까지 소개하는 명령어들이 있습니다. 그리고 극한(limit)을 구하는 명령어들을 모두 모아두었습니다.


    The following website has more information on the materials of the lectures in the 7th week. For contents on multivarable functions, please visit the link.

     http://matrix.skku.ac.kr/Lab-Book/Sage-Lab-Manual-1.htm


    그리고 출력을 얻기 위해서는 print 명령어가 최근 upgrade가 되어, print할 때는 앞뒤로 괄호가 필요합니다. 원하는 도함수(derivative)에 대한 결과를 프린트할 때는 괄호를 주십시오.

    To obtain the output, the print command has recently been upgraded, you need to have a bracket, print ( ), for a proper output.


    그렇게 해서 도함수에 대한 응용이라든지, 앞으로 배울 Newton’s Method, 적분에 대한 코드를 모아놨는데, 필요할 때는 명령어를 활용해서 학습하시면 되겠습니다.


    We have collected codes for Newton’s Method and Integration in the following web address: http://matrix.skku.ac.kr/cal-lab/cal-Newton-method.html


    그리고 마지막으로 제가 미적분학을 가르쳤을 때, 학생들이 미적분학을 쭉 배우면서 자기가 배운 내용들을 쭉 정리한 내용인데요. 참고로 보시면, 학생들이 물어봤던 질문과 저희가 얻은 답변들을 쭉 정리해놨으니까 참고해서 보시면 되겠습니다.


    And lastly, we share PBL reports on the previous calculus class, it's about the record of students who learned calculus class. You can refer to them.


    지금까지 미적분학의 제일 앞부분인 극한, 연속, 도함수, 그리고 접선에 대해서 학습하였습니다.


    So far, we have learned on the limit, continuity, derivative, and tangent line. Those are the very first parts of calculus class.


    [Review] 


    간단히 이번 주 배운 내용을 요약하면, 극한과 도함수에 대해서 이번 7주차에서 학습을 하였습니다.


    In summary, we learned about <Limit and Derivative> in this week.


    7주차 학습의 내용은 <인공지능의 최적해(optimal solution)를 찾기 위해서> 도함수, 접선, 극한값에 대한 개념을 배웠습니다. 특히 함수의 도함수를 극한(limit)을 이용해서 정의했습니다. 그 후 n계 도함수를 구하는 방법을 다루는 미분(differentiation)을 학습했습니다. 이 과정에서 좌극한, 우극한, 수렴, 발산과 점근선, 연속함수, 기울기, 접선에 대해서 학습을 했습니다.


    And we learned the concepts of the derivative, tangent line, and limit. In this process, we learned about left-hand limit, right-hand limit, convergence, divergence, asymptote, continuous function, slope, tangent line. This will help you to find the optimal solution in the problems of artificial intelligence. In particular, the derivative of the function was defined using the limit. In the differentiation, we learned how to obtain the n-th derivative.


    이제 7주 강의를 마치고, 8주에서는 극댓값과 극솟값, 최댓값과 최솟값 개념을 배우고 우리가 인공지능에 필요한 최적해를 구하는 과정에 한 발 더 가까이 가도록 하겠습니다. 수고 많이 하셨습니다.


    In the next class, we will continue to learn the concepts of local maximum, local minimum, maximum value, and minimum value, and take a closer step to the process of finding the best possible solution for artificial intelligence problems. Thank you for your hard work.



        그림입니다.
원본 그림의 이름: KakaoTalk_20200423_152052624.jpg
원본 그림의 크기: 가로 4032pixel, 세로 3024pixel
사진 찍은 날짜: 2020년 04월 23일 오후 3:14
카메라 제조 업체 : samsung
카메라 모델 : SM-A505N
프로그램 이름 : A505NKSU3BTC6
F-스톱 : 1.7
노출 시간 : 1/50초
IOS 감도 : 80
색 대표 : sRGB
노출 모드 : 자동
35mm

     


    Week 8. Introductory Mathematics for Artificial Intelligence

     

                인공지능을 위한 기초수학 입문


    8. 극대, 극소, 최대, 최소  88

    8.1 도함수의 응용

    8.2 2계 도함수의 응용

    8.3 극대, 극소, 최대, 최소


    [8-1 pre]


    안녕하십니까. 이제 8주차입니다. 이번 주에는 ‘극대와 극소, 최대와 최소’에 대해서 학습하겠습니다.


    Hello!! We now in the Week 8. We will study 'Local maximum, Local minimum, Absolute maximum, and Absolute minimum' in this week.


    첫 시간에는 함수의 1계도함수를 이용하여 함수의 증가, 감소를 판단하고, 함수의 2계도함수를 이용하여 함수가 어떤 구간에서 ‘위로 볼록 또는 아래로 볼록’ 인지를 판단하는 방법에 대하여 살펴봅니다. 그리고 이를 이용하여 주어진 함수의 극대, 극소를 판단하며, 최대, 최소를 찾는 원리에 대하여 학습하겠습니다.


    In this first lecture of Week 8, we will learn about how to determine whether a given function is increasing or decreasing on a certain interval by using its first derivative, and how to determine whether the function is 'convex upwards or convex downwards' on an interval using its second derivative. And using this, we will find the <Absolute maximum and Absolute minimum> of a given function on an interval.


    학습 순서는 1절에서는 ‘도함수의 응용’, 2절에서는 ‘2계도함수의 응용’, 3절에서는 ‘극대, 극소, 최대, 최소’에 대해서 학습하겠습니다.


    We will study ‘Application of Derivatives' in <Section 1>, 'Application of Second Derivatives' in <Section 2>, and 'Absolute Maximum and Minimum in <Section 3>.


    기울기 값에 따라서 함수가 한 점 (또는 주어진 구간)에서 증가하는지 감소하는지 도함수를 이용하면 알 수 있습니다. 도함수의 부호가 변하는 과정에서 ‘도함수=0’이 되는 점인 임계점(critical point)를 얻을 수 있습니다. 더 나아가서 1계도함수를 한 번 더 미분한 2계도함수를 이용하면 주어진 함수가 임계점에서 극댓값을 갖는지 또는 극솟값을 갖는지를 알 수 있습니다. 이를 이용하여 함수가 ‘주어진 구간에서 언제 최댓값을 갖는지 또 최솟값을 갖는지를 확인’할 수 있습니다.


    The derivative can be used to determine whether a function increases or decreases depending on the sign of slope at a point (or on a given interval). At the changing point of sign in a derivative (ie. Derivative = 0), we will have a critical point. And the second derivative can be used to determine whether a given function has a local maximum or minimum value at the critical point. Using this, we can 'check when a function has the absolute maximum or minimum on a given interval'.


    이번 주에는 위의 내용에 대해서 학습하도록 하겠습니다.


    Now we start today’s lecture.


    [8주차 1강]  반갑습니다. 이번 8주차 <인공지능과 최적화> 에서는 지난 시간의 극한, 연속, 도함수에 이어서 극대, 극소, 최대, 최소에 대해서 학습하도록 하겠습니다.


    [Week 8, Lec 1] In this eighth week, we will learn about 'local maximum, local minimum, absolute maximum, and absolute minimum' as a part of <AI and Optimization>.


    웹 주소 http://matrix.skku.ac.kr/math4ai-intro/W8/ 에 8주차 교안과 실습실 내용이 있습니다. 학습하면서, 실습실을 활용해서 실제 계산들을 해보시기 바랍니다.


    We can practice the contents at the following link:

    http://matrix.skku.ac.kr/math4ai-intro/W8/


    8장 1절은 <도함수의 응용> 입니다.


    Section 8.1 <Applications of Derivatives>.


    함수 가 구간 에서 정의되어 있을 때, 내의 두 개의 점 , 에 대하여 를 만족하면 는 구간 에서 증가한다고 하며, 인 구간(interval) 내의 임의의 두 점 , 에 대하여 이면 는 구간 에서 감소한다고 합니다. 복잡해 보이지만, 그림에서 쉽게 확인해볼 수 있습니다. 주어진 구간에서 함수값이 항상 증가하면, 함수가 주어진 구간에서 증가한다고 의미이고, 그 구간에서 보다 가 큰데, 보다 항상 더 작으면 그림에서 보듯이 감소한다고 의미입니다.


    Let be a function defined on an interval .

    For any two distinct points and in , if whenever , then we say is increasing on .

    Similarly, if whenever , then we say is decreasing on .

    It will be much easier to understand increasing and decreasing functions from the figure. For example, if f() > f() whenever in an interval, then the function f is a decreasing function on that interval as shown in the figure.


    함수 가 폐구간(closed interval) 에서 연속이고, 개구간(open interval) 에서 미분가능하다고 할 때, 도함수는 증가와 감소에 대한 준거를 제공합니다.


    구간 내의 모든 점에서 이면, 함수 는 에서 증가한다는 의미이고, 구간 내의 모든 점에서 이면, 는 그 구간에서 감소하는 것입니다. 즉 ‘도함수는 주어진 함수가 주어진 구간에서 증가 또는 감소하는 것에 대한 답을 제공한다’는 의미입니다.


    If a function is continuous on a closed interval and differentiable on an open interval , then its derivative may tell us some information on whether f is increasing or decreasing on that interval.


    예를 들어, ( 를 포함하는 구간에서 항상) 이면 에서의 접선의 기울기가 양수이므로, 의 근방에서는 함수가 증가함을 알 수 있다는 의미입니다. 마찬가지로 ( 를 포함하는 구간에서 항상)  이면 의 근방에서는 함수가 감소하는 것을 알 수 있습니다.


    For example, if holds, then the slope of the tangent line to f at is positive, so the function is increasing near . Similarly, holds, then the slope of the tangent line to f at is negative, so the function is decreasing near .


     다음 그림의 이 점 c 에서 기울기가 ‘+’ 즉 양의 방향이니까 점점 증가한다는 의미고, 함수 에서 c점에서의 도함수(derivative)가 음(negative)의 방향으로 있으면, 이 함수는 이 점의 근방에서 감소한다고 보면 됩니다.


    In the following figure, since the slope of the tangent line to f at point c is positive, the function is increasing near c. Similarly, since the slope of the tangent line to f at point c is negative, then the function is decreasing near c.


       그럼 예를 통해서 확인해보도록 하겠습니다.

    Let’s see the following example.


    [예제 1] 함수 가 다음과 같이 주어질 때 함수 가 증가하는 구간과 감소하는 구간을 구해봅시다. (화면의 코드를 보세요)


    [Example 1] Let and find all increasing intervals and a decreasing interval.


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    f(x) = x^3 - 6*x^2 + 9*x + 1

    df(x) = diff ( f(x), x )

    p = plot( f(x), (x, -1, 4), ymin = -5 )  # graph of f

    show ( p )

    print ( "f'(x) = ", df(x) )  # derivative

    solve ( df(x) > 0, x )  # interval where f’(x) > 0

    ------------------------------------------------------------------

                                                             <---화면에 있으니 자막에는 없어도 됨)


    즉, 함수를 그리고, 도함수를 구하고, 도함수가 0보다 크게 되는 x값의 범위를 구하라고 명령어를 주니까, 그래프를 그렸고, f'(x)는 3*x^2-12*x + 9, (x < 1)인 구간과 (x > 3)인 구간에서 증가한다, 그러니까 ‘1까지 증가하고 또 3부터 증가한다’는 답을 줍니다. 따라서 그래프에서 보듯이 그 두 점 사이에서는 함수가 감소한다는 걸 알 수 있습니다. 즉, 증가하는 구간은 (-infinity, 1) 과 (3, +infinity) 이고, 감소하는 구간은 그 사이인 (1, 3)이라는 것을 바로 확인할 수 있습니다.


    This code gives all the needed information. It plots a graph of f(x) = x^3 - 6*x^2 + 9*x + 1, and finds the derivative and an interval where f'(x) > 0. The output is that f is increasing on (-∞, 1) or (3, +∞).

    As shown in the graph, the function is decreasing on (1, 3). So the intervals of increase are (-,1), (3,), and the interval of decrease is (1, 3).


      8.2. Applications of the second derivative (2계 도함수의 응용)


    지금까지는 1계 도함수가 증가구간과 감소구간을 확인하는데 사용되는 것을 배웠습니다. 이어서, 2계 도함수의 응용으로 2계 도함수는 어떤 의미가 있을지, 그 응용에 대해서 확인하겠습니다.


    So far, we have learned that the 1st derivative of a function can be used to identify the increasing or decreasing intervals. Our natural question is what the second derivative would mean and how it can be used? Today we will see an application of the second derivative.


    함수 가 에서 미분가능하고 점 이 곡선 위의 점이라 할 때, 점 를 포함하는 적당한 개구간 가 존재하여 인 구간 의 각 점 에 대응하는 곡선 위의 모든 점 가 점 에서의 곡선의 접선 아래쪽에 있으면, 곡선 는 점 에서 위로 볼록 또는 아래로 오목하다고 하고, 반대로 접선의 위쪽에 있으면 아래로 볼록 또는 위로 오목하다고 합니다. 그림에서 보시면 간단합니다. 위로 쭉 올라가서 도함수가 증가하는 쪽이 되다가 감소하는 쪽이 되면 위로 볼록한 셈입니다. 그리고 쭉 내려가다가 도함수가 음으로 줄어들다가 도함수가 양으로 변하는, 음에서 양으로 변하는 단계가 생길 때는 아래로 볼록하는 그런 변화가 있습니다. 그 의미를 설명한 것입니다.


    Convex functions play an important role in mathematics. They are especially important in the study of optimization problems. Assume that a function is differentiable at and the point is on the curve .  The curve is said to be concave up (or convex down) at if (x, f(x)) lies above the tangent at (c, f(c)) for all x near c. Similarly, the curve is said to be concave down (or convex up) at if (x, f(x)) lies below the tangent at (c, f(c)) for all x near c.

    As shown in the figure, this function goes up (increasing) and reach the top, and then goes down (decreasing), then the function is convex up. We can observe that there was a change of sign on the derivative, from ‘+’ to ‘-'. If the sign of derivative changes from '- to +',  then the function is convex down.


    그러면 기울기가 양에서 음으로 변할 때 위로 볼록, 기울기가 음(negative)에서 양(positive)으로 변할 때, 즉 기울기의 도함수(derivative)가 0 이 되는 그런 변화를 겪을 때, 바로 아래로 볼록한 현상이 생기는 것입니다. 그 말을 풀어쓰면, 곡선 가 점 에서 연속일 때 는 존재하지 않아도 무방하고, 곡선 위의 점 를 경계로 한쪽에서는 위로 볼록하고 다른 쪽에서는 아래로 볼록할 때, 를 곡선 의 변곡점이라고 합니다.


    In other words, a function is convex up in an interval when the sign of the slope of the tangent changes from positive to negative on that interval. A function is convex down in an interval when the sign of the slope of the tangent changes from negative to positive on that interval.


    [Inflection point] If a function is convex up on one side, and convex down on the other side at the point , this point is called the inflection point of the curve . At this point, the existence of does not matter if the curve is continuous at .


    2계 도함수는 함수가 위로 볼록하거나 아래로 볼록한 상황을 알려줍니다. 따라서 함수 가 점 를 포함하는 적당한 개구간 에서 미분가능하고 가 존재할 때, 다음이 성립합니다.


    The second derivative indicates when the function is convex up or convex down. Therefore, when a continuous function is differentiable on an open interval containing the point and does exist, then the followings hold:


     ⓵ 이면, 곡선 는 점 에서 아래로 볼록하다.

     ⓶ 이면, 곡선 는 점 에서 위로 볼록하다.


    ⓵ If , then is convex down at .

    ⓶ If , then is convex up at .


    따라서 볼록하다하는 것은 ‘2계 도함수(second derivative) 가 주어진 구간에서의 0 보다 크냐, 작으냐로 구분’된다는 의미입니다.


    So convexity can be determined by the second derivative of function.


    그래서 이면 이므로 는 의 근방에서 증가함을 알 수 있고, 은 접선의 기울기를 의미하므로, 접선의 기울기가 증가하는 상황입니다.  따라서 다음과 같이 아래로 볼록임을 알 수 있고, 마찬가지로 이면 접선의 기울기가 감소하므로 위로 볼록임을 알 수 있습니다.


    Since and , increases near . And since means the slope of the tangent line, the slope of the tangent line is increasing. Hence, f is convex down at . Similarly, means that the slope of the tangent line is decreasing. Hence, f is convex up at .


    그림에서 분명히 확인하실 수 있습니다.

    You can easily figure out the concept in the .\PICture.

              http://matrix.skku.ac.kr/math4ai-intro/W8/


    [예제 2]는 함수 가 어느 구간에서 위로 볼록하고, 아래로 볼록한지 조사하는 문제입니다. 학습한 대로 함수를 정의하고, 도함수를 구하고, 2계 도함수를 구하고, 함수를 먼저 그려본 후, 2계 도함수를 프린트를 하고 2계 도함수가 언제 0보다 큰지 확인해보겠습니다. (화면의 코드를 보세요)


    [Example 2] Let f(x)= {1} over {3} x ^{3} -x ^{2} -3x+4 . Find intervals where this function is convex up or down.


    그림입니다.
원본 그림의 이름: CLP000034380009.bmp
원본 그림의 크기: 가로 472pixel, 세로 130pixel ---------- http://matrix.skku.ac.kr/KOFAC/ ----------------

    f(x) = 1/3*x^3 - x^2 - 3*x + 4

    df(x) = diff ( f(x), x )  # 1st derivative

    d2f(x) = diff( df(x), x )  # 2nd order derivative

    p = plot ( f(x), (x, -3, 5), ymin = -5 )  # graph

    show ( p )

    print ( "f''(x) = ", d2f(x) )  # 2nd order derivative

    solve ( d2f(x) > 0, x )  # find interval f''(x) > 0

    --------------------------------------------------------------

                                        

    그래서 일 때는 함수는 위로 볼록이고, 구간에서는 함수가 아래로 볼록이다. 부호가 변하는 점이 변곡점인데, 변곡점은 인 이 점이 되는 것입니다.

     

    The output shows that the function is convex up on the interval where , and the function is convex down on the interval where . The changes in the sign was made at point , which is an inflection point.


    간단합니다. 이 내용을 실습해 보겠습니다.


    <인공지능과 최적해> 가 다루는 <미분> 중 <극대, 극소, 최대, 최소> 부분입니다.

    Now consider "local maximum, local minimum, Maximum and Minimum",  which is an important part of "Differentiation".


    도함수의 응용 부분에서 보시면, 증가하는 함수와 감소하는 함수, 도함수들의 크기로 <도함수가 0보다 크면 증가하는 구간>, <0보다 작으면 감소하는 구간> 인 것을 확인할 수 있었습니다.


    As mentioned in <applications of the derivative>, you can see the relationship between the increasing/decreasing functions and the sign of their derivatives. It shows that a function is increasing on an interval where its derivative is greater than zero and a function is decreasing on an interval where its derivative is less than zero.


    이 점 c 에서의 도함수가 0보다 커서 증가하는 접선이 나옵니다. 이 접선으로부터 증가한다는 것을 확인할 수 있습니다. 만일 <점 c 에서의 접선 이 0보다 작으면> 여기서 함수가 감소하고 있는 것을 확인할 수 있습니다.


    The derivative of f at c is greater than zero, so it gives an increased tangent line. It can be seen from this tangent line that function is increasing. If < 0, we can tell that f is decreasing at that point.


     그리고 증가하는 구간과 감소하는 구간을 구하기 위해서 증가하는 구간을 확인해보자고, 함수를 주고 도함수를 구하고 도함수가 0보다 큰 구간을 구하면, 그림을 그리고 도함수를 구해주고 이 부분이 증가하는 구간임을 확인할 수 있습니다.


    If we want to find intervals of increase of a given function f, we find its derivative f’, and then find intervals where f’(x) > 0. From the graph, we can easily see an interval of increase.


     그리고 2계 도함수의 응용에서 설명했듯이, f''(x) =  2*x - 2 = 0이 되는 이 점이 변곡점입니다. 변곡점에서 보듯이 이계도함수의 부호가 변하는 이 점이 변곡점입니다.


    As explained in the application of the second derivative, the inflection point can occur where f''(x) =  2*x - 2 = 0. This point is an inflection point since the sign of the second derivative changes.


    여기서 함수의 기울기가 점점 줄어드는 즉 2계 도함수 의 값이 0보다 작은 이 구간에서 함수는 위로 볼록하게 되고, 2계 도함수가 0인 이 변곡점을 지나서, f''(x) 가 0보다 큰 ‘+ (plus)’가 되는 이 구간에서 아래로 볼록 한 그래프가 됩니다.


    The function is convex up in an interval where the slope of the function is decreasing, i.e. the value of is less than zero, and it is convex down where the second derivative is greater than zero. The point where f''(x)=0 is an inflection point.


    그래서 함수 가 주어졌을 때, 아래로 볼록인지 위로 볼록인지를 구분하기 위해서 ‘함수를 주고 함수의 도함수, 이계도함수를 구하고, 함수의 그래프를 그린 후에 2계 도함수가 0보다 큰 구간이 어디인지를 찾아보라고 명령어’를 주면, 바로 계산해서 함수 그래프를 그려주고, 우리가 원하는 구간, 즉 x>1 구간에서 아래로 볼록한 현상이 나타난다는 것을 보여주는 것입니다. (화면의 코드를 보세요)


    For a given function , the next command can be used to find its first derivative and the second derivative, and then find the interval where the second derivative is greater than zero. It is convex down on x > 1.


    그림입니다.
원본 그림의 이름: CLP000034380009.bmp
원본 그림의 크기: 가로 472pixel, 세로 130pixel ---------- http://matrix.skku.ac.kr/KOFAC/ ---------

    f(x) = 1/3*x^3 - x^2 - 3*x + 4

    df(x) = diff ( f(x), x )  # 도함수

    d2f(x) = diff( df(x), x )  # 2계 도함수

    p = plot ( f(x), (x, -3, 5), ymin = -5 )  # 함수의 그래프

    show ( p )

    print ( "f''(x) = ", d2f(x) )  # 2계 도함수

    solve ( d2f(x) > 0, x )  # 2계 도함수가 0보다 큰 x값의 범위

    ---------------------------------------------------------

                                                             <---화면에 있으니 자막에는 없어도 됨)


    3차가 아니라 4차인 다항식(polynomial)을 주고, y구간을 (-3, 5) 대신  (-15, 5)로 넉넉하게 주면 그래프의 전체모양으로 볼 수 있습니다.


    For a given fourth degree polynomial, we can change the interval and have a wider graph.


    그래프를 보면서 2계 도함수(second derivative)와 2계 도함수가 0보다 큰 x값의 범위를 보면, x가 0보다 작을 때 아래로 볼록이고, x가 2보다 클 때, 변곡점을 지나면서, 아래로 볼록이라는 것을 보여줍니다.


    In this graph, we can see where the second derivative is less than zero, and where it is convex down. When x is greater than 2, the function is convex down after this inflection point.


    이런 방법으로 http://matrix.skku.ac.kr/math4ai-intro/W8/ 주소에서 실습을 할 수 있습니다.

    이번 시간에는 여기까지 하고 다음 시간에 극대, 극소, 최대, 최소를 하겠습니다.  수고하셨습니다.


    We can practice the code at http://matrix.skku.ac.kr/math4ai-intro/W8/.


    In the next session, we will learn how to handle 'local maximum, local minimum, absolute maximum and absolute minimum'. Thank you.


    [8-2 pre]


    8주차 2차시입니다. 우리가 학습 할 ‘극대, 극소, 최대, 최소’ 중, 첫째 시간에는 도함수의 응용으로 ‘1차도함수와 2차도함수’에 대해 학습을 했습니다.


    In the previous lecture, we learned about 'Applications of the first and second derivatives'.


    이번 시간에는 1차도함수와 2차도함수를 이용하여 ‘극댓값, 극솟값, 최댓값, 최솟값’을 구하는 전체 과정을 학습하고 코드를 이용하여 실습하도록 하겠습니다.


    In this second lecture of Week 8, we will study the entire process of finding a local maximum, a local minimum, the absolute maximum, and the absolute minimum by using the first and second derivatives with 'Code'.


    [8주차 2강]  반갑습니다.  계속해서 8주차 2차시에는 앞에서 배운 1계도함수, 2계도함수의 성질을 활용해서 <극대, 극소, 최대, 최소>에 대해서 학습하겠습니다.

               http://matrix.skku.ac.kr/math4ai-intro/W8/

     

    [Week 8, Lec 2] Now we study <local maximum, local minimum, absolute maximum, absolute minimum> by using the properties of the first and the second derivatives we learned earlier. 

    http://matrix.skku.ac.kr/math4ai-intro/W8/


     가 폐구간 에서 연속이면 이 구간에서 가 최댓값을 취하는 점 및 최솟값을 취하는 점이 존재합니다. 이 점을 구하기 위해서는 먼저 극대, 극소에 관하여 알아야 됩니다. 함수 가 의 근방의 모든 점 에 대하여 가 성립하면 함수 는 에서 극댓값 를 갖는다고 하고, 반대로 함수 가 의 근방의 모든 점 에 대하여 가 성립하면 함수 는 에서 극솟값을 갖는다고 합니다. 또, 의 극댓값과 극솟값을 통틀어 극값이라 얘기하며, 극대점 또는 극소점을 극점(extreme point)이라고 합니다.


    If is continuous on a closed interval , then there exists the absolute maximum and absolute minimum of in this interval. To find this, we must know the notion of local maximum and local minimum.

    A function has a local maximum (resp. a local minimum) at if satisfies (resp. ) for all in a neighborhood of . A local maximum or local minimum is called an extreme value, and a point at which a function has a local maximum or local minimum is called an extreme point.


    이 내용을 간단하게 아래 그림에서 확인하도록 하지요.

    이 그래프는 연속이면서 미분가능한 점에서의 극댓값과 극솟값을 보여줌과 동시에, 연속이긴 하지만 도함수가 존재하지 않는 점들도 보여줍니다.

    이제 전체 구간에서의 극댓값, 극솟값들, 즉 극값을 모두 찾았고, 다음에는 전체 구간에서의 최댓값, 최솟값을 찾습니다. 따라서 극댓값들 다 구하고, 양 끝점의 값과 비교해서 그 전체 중에 가장 큰 값이 최댓값이고, 그 중 가장 작은 값이 최솟값이 됩니다.


    Let's see the figure below. This graph shows all local maxima and local minima. The function may have an extreme value at which it is continuous but not differentiable. We found all extreme values in the given domain. After comparing the function values at both endpoints with these extreme values, the largest value among them is going to be the absolute maximum and the smallest value is going to be the absolute minimum of the function on the interval.


    이런 관찰로부터, 극값이 될 수 있는 후보들은 이거나 가 존재하지 않는 점 c 들임을 알 수 있습니다. 이 중간에 있는 값들은 극값이 될 수가 없습니다.  도함수가 0 이거나 도함수(derivative)가 존재하지 않는 점 c 에서만 극값이 생길 수 있습니다.


    From this observation, we see that all possible candidates (to be an extreme point) must be c where or does not exist.


    그래서 함수의 미분계수가 0 이거나 존재하지 않는 점을 그 함수의 임계점(critical point)라고 합니다. critical point 라는 단어가 얘기하듯이 굉장히 중요한(critical) 한 점입니다. 그리고 임계점(critical point)에 대해 다음 정리가 성립합니다.


    A point c where the is zero or does not exist is called a critical point. For each critical point, the following condition holds.


    [Fermat의 임계점 정리]  함수 가 개구간 에서 연속이고 에서 극값을 갖는다면, 이거나 가 존재하지 않는다.



    [Fermat's Theorem for Extrema]

    If f is continuous on an open interval and f(x) has an extreme value at , then or does not exist.

                         [i.e. x must be a critical point.]


    만약에 주어진 개구간에서 함수가 연속이고 극값을 점 c 에서 갖는다면, 이거나 그 점에서 가 존재하지 않는다, 둘 중에 하나입니다.


    이걸 역으로 생각한다면, 우리가 극값을 찾는다면, 우선 이거나 가 존재하지 않는 점들만 찾으면, 극값은 그 중에서만 존재한다고 이해하시면 되겠습니다.


    In other words, all we have to do is first to find all points c where or does not exist.


    [예제 3]는 ‘함수 의 임계점(critical point)을 구하라’는 문제입니다.

    [Example 3] Find all critical points of .


    풀이. 일단 는 다항함수이므로 도함수가 존재하지 않는 점은 생기지 않는다.      따라서 에서 임계점은 또는 이다. (화면의 코드를 보세요)


    Sol) Since f is a polynomial function, there is no point where f is not differentiable. So all critical points are or from the equation  . [see the following code]


    --- http://matrix.skku.ac.kr/KOFAC/ ----------

    f(x) = -2*x^3 + 3*x^2

    df(x) = diff(f(x), x)

    solve(df(x) == 0, x)

    ----------------------------------------------------

    [x == 0, x == 1]                                      ■

                                                           


    결과, 코드를 이용하여 구한 임계점이 손으로 구한 임계점 0과 1과 일치합니다.


    The critical points that we found by using code and by hand are all same.


    이제 도함수를 이용해서 극댓값과 극솟값을 판정할 수 있습니다.


    Now, we can determine all local maxima and local minima.


    함수 가 정의역 내의 한 점 에서 과 을 가지며, 일 때,

     ⓵ 이면, 는 함수 의 극댓값입니다. [는 임계점 에서 극댓값을 갖습니다]

     ⓶ 이면, 는 함수 의 극솟값입니다.


    When a function has and at c with ,

    ⓵ if , then is a local maximum.

    ⓶ if , then is a local minimum.


    이 그림은 임계점에서 극댓값 또는 극솟값을 갖는 것을 보여줍니다.

    그래서 이면 점 에서 위로 볼록인데, 이므로 아래 그림에서 는 함수 의 극댓값임을 쉽게 알 수 있다. 마찬가지로 이면 아래로 볼록이므로 다음 그림에서 는 함수 의 극솟값임을 쉽게 알 수 있습니다.


    This figure shows that a function has a local maximum or local minimum at a critical point. If , then is a convex up at the point . And implies that is a local maximum. Similarly, the is a local minimum of if .


    폐구간에서 연속인 함수 의 최댓값과 최솟값은 임계점에서의 함숫값과 구간의 양 끝점에서의 함숫값을 비교하여 구하면 된다는 말입니다. 즉


    [단계 1]  구간 에서 의 임계점들을 찾는다.

    [단계 2] 이 각각의 임계점들과 양 끝점에서 의 값을 계산한다. 그 중 가장 큰 값이 최댓값이고, 가장 작은 값이 최솟값입니다.

     

    The absolute maximum and absolute minimum of a continuous function in a closed interval can be obtained by comparing the function values at all critical points and the function values at both ends of the interval.


    [Step 1] Find all critical points of on interval .

    [Step 2] Compute function values at all critical points and end points.


    The largest value among them is the absolute maximum, and the smallest value is the absolute minimum of the function on I.


    [예제 4]는 ‘ 함수 가 구간 에서 정의되었다고 할 때, 의 극댓값과 극솟값을 구하여라. 또, 의 최댓값과 최솟값을 구하시오.’ 라는 문제입니다. 문제를 풀기 위해, 먼저 =0이라고 놓고 critical point들을 찾고 그 값에서 함숫값들을 찾은 후에 양쪽 끝점, x=-2인 점과 x=6인 점과 다 비교해서 그 중에 가장 큰 값, 가장 작은 값을 찾으면 되겠습니다. (화면의 코드를 보세요)


    [Example 4] Find all local maxima and local minima of .

     is defined on interval .

                          (read the code in the screen)


    ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    f(x) = 1/3*x^3 - x^2 - 3*x + 4

    df(x) = diff(f(x), x)

    d2f(x) = diff(df(x), x)

    solve(diff(f(x)) == 0, x)  # critical points [x == 3, x == -1]

            # 임계점 [x == 3, x == -1] 을 f 위 2계도함수에 대입하면 

    print(d2f(-1), d2f(3))  # x= -1, local maximum, x = 3, local minimum

    print("local maximum =", f(-1))

    print("local minimum =", f(3))

    print(f(-2), f(6))  # function values at end points.

    ----------------------------------------------------------------

    -4          4

    local maximum =   17/3

    local minimum =    -5

    10/3       22    # 구간의 양 끝점에서의 함숫값


                                                             <---화면에 있으니 자막에는 없어도 됨)


    위에서 보듯이 주어진 구간에서 의 극댓값은 , 극솟값은 , 최댓값은 , 최솟값은 입니다.  


    즉 임계점(critical point) 2개를 찾아서, 그 점에서의 값, 극값 2개와 양 끝점에서의 값을 비교하여, 4개 중에서 최댓값과 최솟값을 구했습니다. ■


    We found two critical points, and we computed the function values at those points and both endpoints. These four values were compared to find the absolute maximum and absolute minimum of the given function. The answer is as follows:

    local maximum :

    local minimum :

    absolute maximum:

    absolute minimum :


    [열린문제 2]  여러분들이 다른 책에서 아주 복잡한 두 번 미분가능한 함수 (또는 두 번 이상 미분가능한 함수)를 찾아서 그 함수의 극댓값과 극솟값 그리고 그 구간에서의 최댓값과 최솟값을 같은 코드를 이용해서 구해보시면 되겠습니다.


    [Open problem 2] Find a complicated twice differentiable function in other textbooks, and use the code to find the local maximum, local minimum, absolute maximum, and absolute minimum of the function.


    이어서 [Fermat의 임계점 정리]에 의해 최적해(optimal solution) 는 다음을 만족합니다.

                                   

    따라서 방정식 을 풀어서 나온 해들이 최적해가 되는지 판단하면 됩니다. 그러나 함수 가 복잡한 경우는 방정식을 풀어서 임계점을 구하는 것조차도 쉽지 않습니다. 이런 경우에는 수치적인 방법으로 임계점을 구하면 됩니다. 대표적인 수치 최적화 방법인 경사하강법(gradient descent method)에 대하여는 다음 절에서 소개하겠습니다.


    By Fermat's Theorem for Extrema, the optimal solution satisfies:

                                  .


    Therefore, we solve an equation , and determine which critical point gives us an optimal solution. However, if the function is too complicated, it is also difficult to find all critical points from the equation . In such cases, we can find them by using a numerical method. The gradient descent method is one of numerical optimization methods for solving this problem. We will study the gradient descent method in the next class.


    오늘은 여기까지 학습하고 실습을 하겠습니다. 임계점을 찾는 방법인 경사하강법(gradient descent method)에 대해서는 바로 다음 절에서 설명하도록 하겠습니다.


    여기서 극값에 대해 더 알고 싶으면 미적분학의 실습실  [미적분학 실습실] http://matrix.skku.ac.kr/Cal-Book1/Ch4/에서 확인하시고, 연습문제 풀이들은  http://matrix.skku.ac.kr/Cal-Book/part1/CS-Sec-4-1-Sol.html 에서 볼 수 있습니다.


    For m
    ore information, we can refer to the following link:

    http://matrix.skku.ac.kr/Cal-Book1/Ch4/

    and solutions of problem can be found in

    http://matrix.skku.ac.kr/Cal-Book/part1/CS-Sec-4-1-Sol.html


    이제 극대, 극소, 최대, 최소에 대해서 실습을 합니다.


    Now we do practice the contents on local maximum, local minimum, absolute maximum and absolute minimum.


    ‘함수의 임계점을 구해라’고 질문하면, 도함수를 구해서 이라고 놓고 풀면 됩니다. 그래서 주어진 함수를 주고, = 0이라고 놓아서 해가 0과 1인걸 알 수 있습니다. 다른 함수에 대하여도 어려움 없이 같은 코드를 이용해서 근을 구할 수가 있습니다.


    First, find all critical points of a function. It can be done easily by using code.


    이 되는 값을 항상 구할 수 있습니다. 그리고 그 함수의 1계도함수 뿐만 아니라 2계 도함수도 구할 수 있습니다. 이때, 1계 도함수가 함수의 증가, 감소를 의미하며, 2계 도함수는 기울기의 증가, 감소를 의미한다는 것을 배웠습니다.


    It is always possible to obtain a value of and . The first derivative tells us where the function is increasing or decreasing. And the second derivative tells us where the 1st derivative of the function is increasing or decreasing.


    그래서 실제 실습을 해보면, 함수가 주어졌을 때, 극댓값, 극솟값, 최댓값, 최솟값을 구하려면, 함수를 정의하고, 1계도함수, 2계도함수를 구해서 방정식 풀고, 극값 구하고, 양쪽 끝점에서의 함숫값을 구해서, 그 값들 중 가장 큰 값 22가 최댓값이 되고, 그리고 가장 작은 값인 –5가 최솟값이 되는 것도 확인했습니다.


    For a given function, to obtain a local maximum, local minimum, absolute maximum and absolute minimum, all we have to do is to define a function, find its 1st derivative, the second derivative, and critical points by solving the equation f’(x) = 0, and then compare the function values at critical points and endpoints.

    In this case, the absolute maximum of the function on the given interval is 22, and the absolute minimum is –5.


    다음 시간에는 경사하강법에 대하여 학습하겠습니다.


    In the next class, we will learn about <gradient descent method> which is crucial in machine learning.


    [Review]


    이번 주에 배운 내용을 요약해 봅시다.


    극대, 극소, 최대, 최소에 대해서 학습을 했으며, 극댓값, 극솟값을 판단하기 위해서 1계도함수, 2계도함수를 어떻게 활용하는 지를 배웠습니다.


    We learned about a local maximum, local minimum, absolute maximum, and absolute minimum, and how to use the derivative, the second derivative to determine the local maximum and local minimum.


    순서는 도함수의 응용, 2계도함수의 응용으로 아래로 볼록, 위로 볼록을 배웠습니다. 그리고 극솟값, 극댓값, 과 양쪽 끝점에서 값을 비교하여 최댓값과 최솟값을 구하는 방법을 학습을 한 것입니다.


    We learned about the application of the derivative (interval of increase and decrease), the application of the second derivative, and convexity (convex down and the convex up). We also learned how to obtain the absolute maximum and minimum by comparing all extreme values and the function values at both endpoints.


    구체적으로는 기울기를 이용하여 함수의 증가와 감소를 비교하고, 2계도함수를 이용하여 함수가 아래로 볼록한지 위로 볼록한지를 판단하였습니다.  즉 인 c를 찾아서 만일 이면 극댓값, 이면서 이면 극솟값을 가진다는 것을 배웠습니다.


    We learned how to determine the convexity of given functions by using the second derivative of the function and how to determine a local maximum or local minimum. It was done by finding all c's that satisfy , then we can tell f(c) is the local maximum if , and f(c) is the local minimum if .


    수고하셨습니다. 다음 시간에는 미적분학에서 배우는 내용 중 <인공지능에서 가장 중요한 지식>인  <경사하강법> 에 대해서 학습하도록 하겠습니다.  감사합니다.


    In the next class, we will learn about the most important algorithm in Machine Leaning, which is known as 'the gradient descent method'. Thank you.

           그림입니다.
원본 그림의 이름: 스트롱코리아포럼2020 (준) 세셔2 62.jpg
원본 그림의 크기: 가로 4673pixel, 세로 2414pixel
사진 찍은 날짜: 2020년 05월 27일 오후 13:10
카메라 제조 업체 : NIKON CORPORATION
카메라 모델 : NIKON D4S
프로그램 이름 : ViewNX 2.10 W
F-스톱 : 6.3
노출 시간 : 1/160초
IOS 감도 : 1600
노출 모드 : 수동
35mm 초점

     

    Week 9. Introductory Mathematics for Artificial Intelligence

     

                인공지능을 위한 기초수학 입문

     


    9. 경사하강법, 최소제곱문제의 해 95

    9.1 경사하강법

    *9.2 응용(최소제곱문제)

     - 과제 (열린문제)-  107


    [9-1 pre]


    안녕하십니까. 9주차입니다. 이번주에는 경사하강법(gradient descent method)와 경사하강법을 이용한 최소제곱문제의 해법에 대해서 학습하도록 하겠습니다. 함수의 최적해를 구하는 대표적인 수치 최적화 방법인 경사하강법을 학습하는 것으로 시작합니다.


    Hello. In Lecture 9, we will learn about the gradient descent method and its application to the least squares problem. The gradient descent method is a simple and most important numerical optimization algorithm for minimizing a function.


    We now start the gradient descent method which is the most popular numerical optimization method to find the optimal solution in AI.


    주어진 함수의 최적해를 계산하려면 도함수가 0이 되는 임계점(critical point)을 구한 후 극대, 극소, 최대, 최소가 되는지를 판단해야 합니다.


    To find an optimal solution for a given minimization problem, we require to compute the critical point at which the derivative is zero, and then determine whether the function has a local maximum, local minimum, the absolute maximum, or absolute minimum at that point.


    그러나 함수가 복잡한 경우에는 임계점을 구하는 것조차도 쉽지 않습니다.

    However, if the function is too complicated, it is not even easy to find critical points.


    이라는 방정식을 풀어서 해를 구하는 문제는 일반적으로 쉽지 않은 문제입니다. 이 때 사용되는 대표적인 방법으로 경사하강법이 있습니다. 경사하강법은 딥러닝의 가중치를 업데이트 하는데 사용되는 핵심알고리즘이기도 합니다. 이번 시간에는 이 경사하강법에 관하여 자세히 살펴보겠습니다.


    In this case, we use the gradient descent method to solve . Gradient descent method is also the core algorithm used to update the weights in deep learning. Let's take a closer look at this gradient descent method.


    1차시에는 경사하강법에 대해서 학습하고, 2차시에는 경사하강법을 이용하여 최소제곱문제 해를 구하는 과정을 학습합니다.

    Today we learn the gradient descent method, and in the second class, we will learn how this gradient descent method can be used in the process of solving the least squares problem.


    경사하강법 알고리즘은 다음과 같이 3단계로 나누어지고 그 과정은 의외로 간단합니다. 저희가 한 단계에서 를 구하면, 거기에 학습률 에타()와 방향벡터 를 곱한 것을 빼줘서 다음 인 로 정해주는 과정을 반복하는 내용으로 함수의 극값 근처에서 시작하여 단계를 거치면서 점점 의 변함에 따라 가 수렴하는 극한값을 찾는 과정입니다.


    The gradient descent algorithm is surprisingly simple. The iteration form is

    . As iterations proceed, the function value of f(x) is updated accordingly and converges to a local minimum.



    이때, 우리가 도함수를 활용하고 접선을 이용하며, 1변수 함수가 아니라 2변수 함수, 3변수 함수, 다변수 함수일 경우에는 도함수의 일반적인 개념인 그래디언트(gradient)를 이용하여 최소제곱문제를 풀어가고 경사하강법은 의 극값 근처에서 시작하여 알고리즘을 거치면서 점점 극값에서의 함수값을 주는 방향으로 수렴하게 됩니다.


    For minimizing a given function of several variables, we use its gradient, which is the generalization of the derivative of a one variable function, instead the derivative. From the gradient descent algorithm, the function values of iterates converge to a local minimum.


    인공지능, 머신러닝의 핵심 알고리즘인 ‘경사하강법’에 대해서 자신감을 갖고 설명할 수 있도록 이번 시간에 모두 집중하시기를 기대합니다.


    We look forward to see your attention today on this ‘gradient descent method’ which is a core algorithm of AI and machine learning. So you can explain later with full of confidence.

    [9주차 1강]  http://matrix.skku.ac.kr/math4ai-intro/W9/


    안녕하십니까. 9주차 K-MOOC [인공지능을 위한 기초수학 입문]에 가장 중요한 주제의 하나인 <경사하강법>에 대해서 학습하도록 하겠습니다.


    [Week 9, Lec 1] Today, we will learn one of the most important part in this class, 'the gradient descent method(GDM)'.


    9주차 강의의 제목은 <경사하강법과 최소제곱문제의 해>입니다. 이 경사하강법을 이용해서 최소제곱문제의 해를 구하는 방법을 학습하도록 하겠습니다.


    In this class, we will introduce 'the gradient descent method(GDM)' and explore 'how to use GDM to solve the least square problem'.


    경사하강법(Gradient Descent Algorithm)에 대해서는 제가 만든 웹 주소에서 실습해보실 수 있습니다.


    You can practice the Gradient Descent Algorithm at the following link:

    http://matrix.skku.ac.kr/math4ai/gradient_descent/ 


    최소제곱문제의 해를 행렬을 이용해서 구하는 방법, QR분해를 이용해서 구하는 방법은 앞에 선형대수학을 배울 때 학습했습니다. 이번에는 연속함수에 대해서 수치적으로 근사해를 얻는 경사하강법을 소개하겠습니다.


    In the previous lecture, we discussed the method of finding the optimal solution of a least squares problem by using QR decomposition. Today we introduce the gradient descent method for solving the same problem numerically.


    다음과 같이 미분 가능한 일변수 함수 의 최솟값을 구하는 문제

                                      

    가 있다고 하겠습니다. 여기서 의 그래프가 다음과 같은 그림 모양입니다. 이 점에서 극솟값이 존재합니다. 이 값을 어떻게 효과적으로 찾을 수 있을까요? 앞에서 배운 미분, 도함수, 이차도함수, 볼록, 오목 같은 성질들을 활용하도록 하려고 합니다.


    We consider the minimization problem , where is differentiable and the graph of f is as shown in the figure below. How can we find this local minimum (also absolute minimum) effectively? One answer is to use the knowledge of calculus, including derivative and convexity which we have learned earlier.


    그러면 앞서 학습한 [Fermat의 임계점(critical point) 정리]에 의해서, 만약에 이 함수가 최솟값을 갖는다면, 이 최솟값에서는 이 되어야 하기 때문에 먼저 함수의 최솟값을 주는 , 우리는 이것을 최적해(optimal solution) 이라고 부릅니다. 이 는 꼭 이 식 을 만족해야만 하는 것입니다.


    By [Fermat's Theorem for Extrema], if this differentiable function has the local minimum value at , then should be hold at .


    따라서 방정식 를 만족하는 해들을 먼저 구해서 그 중에 어느 것이 최적해가 되는지 판단하면 됩니다. 그러나 함수 가 복잡한 경우는 방정식을 풀어서 임계점을 구하는 것조차도 쉽지 않습니다. 이런 경우에는 수치적인 방법으로 임계점을 구합니다.


    Therefore, we have to find all solutions that satisfy the equation , and then determine which of them is going to be the optimal solution. However, when the function is too complicated, it is not easy to solve the equation for finding critical points. In such cases, the critical point can be obtained by a numerical method.


    이 절에서는 주어진 함수의 최솟값을 구하는 대표적인 수치 최적화 방법인 경사하강법(GDM)에 대하여 살펴봅니다. 경사하강법의 기본 아이디어는 함수의 기울기(경사)를 구하여 기울기가 낮은 쪽으로 계속 이동시켜서 극값에 이를 때까지 반복시키는 것입니다. [경사하강법] 알고리즘을 요약하면 다음과 같습니다.  (화면의 표를 보세요)


    Gradient Descent Method(GDM) is the simplest numerical optimization method for finding the minimum value of a given function. The basic idea of GDM is iteratively moving in the direction of steepest descent as defined by the negative of the gradient. The algorithm is summarized as follows.

     


    [경사하강법] 알고리즘 [Gradient descent algorithm]

     

    [단계 1]  초기 근사해 , 허용오차(tolerance) , 학습률(learning  rate) (eta, 에타)를 준다. 이라 한다.

    [단계 2]  를 계산한다. 만일 이면, 알고리즘을 멈춘다.

    [단계 3]  , 이라 두고 [단계 2]로 이동한다.


    [Step 1] Set an initial iterate , tolerance , initial learning rate (eta) and iteration number .

    [Step 2] Compute . If , then stop.

    [Step 3] Set , and go to [Step 2].


    여기서 은 인데, 부등호 두 개 << 를 쓴 의미는 1보다 아주 작다는 의미입니다. 그러니까 0에 아주 가까운 작은 (non-negative) 을 허용오차와 학습률(learning rate) (eta, 에타)를 주고, 를 1로 정의하고 시작한다는 의미입니다. 단계를 거쳐서 의 값이 우리가 정한 허용오차 이내에 들어오면, 알고리즘을 멈추는 것입니다.


    Here, epsilon (tolerance) satisfies . A notation ‘<<’ in this inequality means that this epsilon is much smaller than 1. Hence, we give a small tolerance epsilon, which is very close to zero, and with a learning rate , then let the iteration number k be 1. If the value of is within a given tolerance after some repeated step , then the algorithm stops.


    먼저 알고리즘의 작동방식을 살펴봅니다. 우선 임의로 을 정하여 을 계산합니다. 만일 (여기서 (epsilon, 엡실론)은 과 같이 매우 작은 양수를 의미합니다.) 이 성립하면  이 되어 (우리가 허용하는 오차범위에서) 임계점 정리를 만족하므로 알고리즘을 멈추고 을 최적인 근사해로 제공합니다. 그러나 이면 계산공식 을 이용하여 를 결정합니다. 같은 방법으로 를 계산하여 임계점 정리를 만족하는지 판단해보고, 만일 성립되지 않으면 를 결정합니다. 이런 방식으로 , , ()를 만들어 냅니다. 마침내 을 만족하는 극한값 에 근사한 를 답으로 제공하고 알고리즘은 마칩니다.

     

    Let's look at the algorithm.

    1) Choose an arbitrary and compute .

    2) If , then it means and critical point theorem is approximately satisfied at , so the algorithm stops and this is an optimal solution. (Here means a very small positive number, for example .)

    3) However, if , compute and go back to step 1) with .

    In this way, we will get ,,... and finally which is very close to the extreme value that satisfies ,and stop the algorithm.


    이제 경사하강법의 원리를 자세히 살펴보겠습니다. 아래 그림과 같이 번째 단계의 근사해 에서의 접선의 기울기가 을 만족한다고 합시다. 그러면 에서 오른쪽으로 이동할 때, 양의 방향 함수가 감소하므로 최솟값을 갖는 최적해 는 의 오른쪽에 있다고 판단할 수 있습니다. 따라서 에서 방향(오른쪽)으로 이동하여 을 생성하게 되는 것입니다. 즉, 가 있으면 이 오른쪽으로 옮겨가게. 이 과정을 조금 조금씩 반복하면서, 가 0으로 수렴하게 만드는 것입니다.


    Let's see in detail how the GDM works. Suppose the slope of the tangent line to the function is negative at the approximate solution at the -th iteration, that is, as shown in the figure below. This means the function is decreasing when moves from left to right, so we can expect x* is located in the right side of x_k. If x_k moves from the left to right, then x_{k+1} is more closer to the optimal value x*. Therefore, we move from to with direction.


     마찬가지로 아래 그림과 같이 번째 단계의 근사해 에서의 접선의 기울기가 을 만족한다고 합시다. 그러면 에서 왼쪽으로 이동할 때(음의 방향) 함수가 감소하므로 최솟값을 갖는 최적해 는 의 왼쪽에 있다고 판단할 수 있습니다. 따라서 에서 방향(왼쪽)으로 이동하여 을 생성한다. 이런 방식으로 , , ()를 만들어 냅니다. 마침내 을 만족하는 극한값 에 근사한 를 답으로 제공하고 알고리즘은 마칩니다.


    Similarly, the slope of the tangent line to the function is positive at the approximate solution at the -th iteration, that is, as shown in the figure below. This means the function is increasing when moves from left to right, so we can expect x^* is located in the left side of x_k. If x_k moves from right to left, then x_{k+1} is more closer to the optimal value . Therefore, we move from to with direction. If we repeat the process until moves to , then will converges to 0. In this way, we get the approximate solution .


        어느 방향에서 시작하여도  인 로 수렴하게 되는 것입니다.


    So no matter which direction we start, converges to such that with this algorithm.



    이를 통해 경사하강법은 가 만족되도록 , , , 을 찾으려고 하는 것임을 알 수 있습니다. 이때 얼마만큼을 이동해야 하는지는 (에타)가 결정하는데, 이를 학습률(learning rate)이라고 합니다. 학습률이 너무 크면 를 넘어서 지나칠 수 있고, 심지어 함수값이 증가하여 수렴하지 않을 수도 있습니다(아래 왼쪽 그림).


    This shows that the GDM will find , , , that satisfy . Here, determines a magnitude of movements. We call this the ‘learning rate.’ If the learning rate is too large, can pass over , or even the function value may increase. So we should be careful to choose a reasonable learning rate .


      반대로 너무 작으면 수렴하는 속도가 느릴 수 있습니다(아래 오른쪽 그림). 적절한 를 이용하여 계산공식 에 따라 을 결정하는 것이 key 아이디어 입니다. (너무 빨리 서두르면 수렴하지 않게 되고, 너무 또 안전하게 하다 보면 시간이 오래 걸리고 수렴속도가 길어지는 경우가 있으니까 를 현명하게 잡아주어야 합니다.)


    On the contrary, if this learning rate is too small, then the rate of convergence will be very slow. So it is important to use proper and a reasonable learning rate .


    [참고]

    적절한 학습률을 결정하는 것은 경사하강법에서 중요한 문제이지만 여기서는 자세히 다루지 않습니다. 학습률은 대개 에서 사이의 범위에서 정하는 것으로 알려져 있으며, 초기 학습률로는 또는 이 주로 사용됩니다.


    [Note] It is important to determine the appropriate learning rate, but we leave the detail of this issue for the next stage. It is usual to set in between and . As an initial learning rate, or is usually chosen.


    예제 5. 함수 의 최솟값을 구하시오. 단 , , 으로 한다.

    풀이. 에서 임계점은 이고 는 아래로 볼록인 이차함수이므로 에서 최솟값 임을 알 수 있습니다.


    Example 5. Find a minimum of .

    Take, ,, and in the following code.


       함수 의 그래프를 그리면 아래로 볼록인 이차함수이고, 임계점을 찾았으니까 그 점에서 최솟값을 갖는 것을 알 수 있습니다. 이번에는 경사하강법(Gradient Descent Algorithm)에 대한 코드를 활용해서 최솟값이 나오는지 한번 확인해보도록 하겠습니다. (화면의 코드를 보세요)


    ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    f(x) = 2*x^2 - 3*x + 2  # function

    df(x) = diff(f(x), x)  # derivative

    x0 = 0.0  # initial iterate

    tol = 1e-6  # tolerance

    eta = 0.1  # learning rate


    for k in range(300):

        

        g0 = df(x0)

        

        if abs(g0) <= tol:

            

            print("Algorithm Success!")

            break

        

        x0 = x0 - eta*g0


    print("x* =", x0)

    print("|g*| =", abs(g0))

    print("f(x*) =", f(x0))

    print("iteration number =", k + 1)

    ------------------------------------------------------------------

    Algorithm Success!

    x* = 0.749999834194560

    |g*| = 6.63221759289456e-7

    f(x*) = 0.875000000000055

    iteration number = 31                            ■

                                                        


    31회의 iteration 만에 임계점 에서 최솟값 을 구해줬습니다. 이렇게 경사하강법(Gradient Descent Algorithm)이 작동하는 것을 확인했습니다. 다음 함수 의 최솟값도 구해보실까요.


    The minimum was obtained at the critical point after 31 iterations.


    We could use the code of the GDM for the next problem. We only need to change the function in the above code. We can get the new answer right away.


    [열린문제 3] 함수 의 최솟값을 구하시오. 단 , , 으로 한다.


    코드에서 함수만 바꿔주면 조건은 모두 똑같으니까 답을 바로 구하실 수가 있습니다.


    [Open problem 3] Find a minimum value of .

    Here, use , , and with the above GDM code.


     지금까지 경사하강법의 알고리즘과 작동 원리를 간단한 예제를 들어 설명하였습니다. 경사하강법은 딥러닝에서 가중치를 업데이트하는 데 사용되는 핵심 알고리즘입니다.  앞서 경사하강법을 소개할 때 정의역 전체에서 아래로 볼록(convex)인 함수를 가정하였습니다. 그러나 구간마다 아래로 볼록, 위로 볼록이 모두 포함된 함수(non-convex, 예. 아래 그림)에 경사하강법을 적용하면, 시작점 이 어디냐에 따라 서로 다른 극솟값 또는 (극대도 극소도 아닌) 임계점으로 수렴할 수도 있다는 것을 염두에 두셔야 됩니다. 


    So far, the algorithm and the principle of GDM have been explained with a simple example. The GDM is a key algorithm that are used to update weights in Deep Learning. When we introduce the GDM algorithm, we assumed a function that is convex down on domain. However if we have a function that contains both convex down and convex up in an interval (non-convex), this GDM algorithm may not work well as you ca see in the figure below. in that case the algorithm may converge to different points (other local minimum, some may not even a extreme or critical points) depending on where the starting point is given.


    [열린문제 4] 위의 그림과 같이 다양한 미분가능 함수인 경우, 일단 그래프의 개형을 그리고, 눈으로 확인되는 극솟값을 포함하는 작은 구간들을 정한 후, 각각의 구간에서 시작점을 잡아 경사하강법을 적용합니다. 그러면 큰 무리 없이 적용됩니다. 이 때 얻은 결과값(output)에 대하여 토론해 보십시오.


    [Open problem 4] For various differentiable functions, as shown in the figure above, draw a graph. Determine a small interval containing the local minimum identified from the figure (by your eye), and apply the GDM code to each interval by take a reasonable starting point .


    그래서 우리의 첫 강의에서 <함수 그래프> 그리는 실습을 한 것입니다.


    이제 경사하강법에 대한 중요한 내용은 모두 배웠습니다. 그러나 경사하강법 이외에도 최적해를 구하는 Newton의 방법이 있습니다. 이는 고등학교 수준을 넘습니다. Newton의 방법과 테일러전개 등에 대한 내용과 함께 궁금한 사항은 아래를 참고하십시오.    [Newton의 방법] http://matrix.skku.ac.kr/Cal-Book1/Ch4/

    [인공지능과 최적해] http://matrix.skku.ac.kr/math4ai/gradient_descent/


    We learned about the GDM. In addition to that, there is a Newton's method of obtaining the optimal solution. This is above the high school level mathematics. Visit the following link for more information on Newton's method.


    [Newton’s Method]  http://matrix.skku.ac.kr/Cal-Book1/Ch4/

    [AI and optimal solution]

              http://matrix.skku.ac.kr/math4ai/gradient_descent/



    여기서 마치고 이어서 9주차 2차시에는 <최소제곱문제>과 <경사하강법의 응용>을 학습하도록 하겠습니다. 수고 많이 하셨습니다.


    Now we close up today's lecture, and we will continue to study "Application of GDM to the Least Square Problem" later. Thank you.


    [9-2 pre]


    안녕하십니까. 지난 시간에는 ‘경사하강법’이 작동하는 원리에 대하여 학습하였습니다.


    Hello. In the previous session, we learned how the ‘gradient descent method’ works.



    이번 시간에는 ‘경사하강법의 응용’으로 앞에서 배운 ‘경사하강법을 최소제곱문제에 적용하여 최적해를 찾는 과정’을 학습하겠습니다.

               http://matrix.skku.ac.kr/math4ai-intro/W9/


    Today in the second session of <Application of the gradient descent method>, we will study 'a process of finding the optimal solution to a least square problem'  by applying this gradient descent method.

               http://matrix.skku.ac.kr/math4ai-intro/W9/

    *[9주차 2강] http://matrix.skku.ac.kr/math4ai-intro/W9/


    안녕하십니까. 9주차 1차시에서 우리는 경사하강법(Gradient Descent method)에 대해 배웠습니다.

     

    [Week 9, Lec 2] In the first session of this week, we have learned the gradient descent method(GDM).


       2차시 <경사하강법의 응용>에서는 <경사하강법을 이용해서 최소제곱해를 구하는 문제>를 학습하겠습니다. (어렵다고 느끼시면 이 절은 skip 하셔도 됩니다.)


    In the second session, under the to.\PIC on <The application of the gradient descent method(GDM)>, we will go through a problem on 'Finding the least square problem by using GDM'. You may skip this section if you feel this part is too difficult.


        우리는 라는 연립방정식을 풀기 위한 최소제곱해법을 앞에서 배웠습니다. 지금은 라는 선형연립방정식이 아니라 연속함수에 대한 최소제곱법 문제

                       min

    를 풀려고 합니다. 이 최소제곱문제는 오차(error)를 최소화시키는 그런 해(solution) u를 구하는 문제입니다. 이 문제를 경사하강법으로 해결할 수 있습니다.


    In the previous lecture, we learned the least square solution to an inconsistent system of linear equations . A least square solution solves the equation as closely as possible, in the sense that the is minimized. Now we try to solve this least squares problem by using the gradient descent method.


    앞서 소개한 경사하강법은 독립변수가 하나인 일변수 함수 에 대한 것입니다. 여기서는 독립변수가 최소 2개 이상인 다변수 함수로 다음과 같이 일반화하여 학습합니다. (화면의 표를 보세요)


    The GDM introduced in the first session is for minimizing an one variable function . Let's generalize the algorithm for minimizing a function of several variables.


    [다변수 함수에 대한 경사하강법] 알고리즘

    [단계 1]  초기 근사해 , 허용오차(tolerance) , 학습률(learning

             rate) 를 준다. 이라 한다.

    [단계 2]  를 계산한다. 만일 이면, 알고리즘을 멈춘다.

    [단계 3]  , 이라 두고 [단계 2]로 이동한다.

                                                          

    [GDM Algorithm for minimizing a multi-variable function]

    [Step 1] Set an initial iterate , tolerance , initial learning rate (eta), and iteration number .

    [Step 2] Compute . If , then stop.

    [Step 3] Set , and go to [Step2].


    앞에서 본 것과 똑같은 알고리즘입니다. 유일한 차이는 변수가 1개에서 n개로 늘어난 것이고, 도함수가 그래디언트(gradient) 로 바뀐 것뿐입니다.


    Basically, it's the same algorithm as the GDM we have learned in the first session. The differences between these two algorithms were the number of variables and the derivative was replaced by the gradient.


    다변수 함수에 대한 것도 기본적인 구성은 일변수 함수의 경우와 완전히 같습니다. 스칼라 가 벡터 로, 절댓값 이 벡터의 노름 으로, 도함수 가 다변수함수에 대한 도함수 역할을 하는 그래디언트(gradient) 로 바뀌는 것만 차이가 있습니다. 다변수 함수와 그래디언트에 대한 개념은 대학수학에서 학습하니까, 또는 이미 배웠을 테니까, 여기서는 기본적인 정의만 아래에 소개하고 실제로 구할 수 있도록 소개해드리겠습니다.


    All the steps of GDM algorithm for minimizing a multi-variable function are exactly same as for one variable function.


    ■ One variable function ⟷ Multi-variable function

    ① Scalar ⟷  Vector

    ② Absolute value ⟷  Norm of vector

    ③ Derivative ⟷ Gradient .


    Since the definition of a multi-variable function and its gradient are learned in college mathematics, let's review them again and learn how to actually find them.


      를 독립적으로 변화하는 두 변수라 하고 를 제 3의 변수라 한 후에 의 값이 각각 정해지면 여기에 대응하여 의 값이 정해질 때 로 표시합니다. 같은 방식으로 더 많은 변수의 함수도 정의할 수 있습니다. 이와 같이 독립변수가 여러 개 있는 함수를 다변수함수라고 합니다.


    Let x and y be independent variables and z be a third variable. Now we can consider a function , which is depend on x and y. Functions of several variables can be defined in the same manner.


      좌표 공간에서 및 는 또 하나의 점 로 생각할 수 있고 와 가 움직이면 일반적으로 하나의 곡면이 움직이게 돼서 의 그래프는 곡면이 됩니다. 예를 들어, 적당한 범위에서 2 변수 함수 의  (3차원) 그래프를 그리면 다음과 같습니다. (화면의 코드를 보세요)


    In a 3-dimensional (coordinate) space, can be considered as a point . As and move, the graph of becomes a surface. For example, using the codes, the graph of a two variables function can be drawn as follows.


    -----------------------------------------------------

    var('x, y')  # variables

    f(x, y) = -x*y*exp(-x^2 - y^2)

    plot3d(f(x, y), (x, -2, 2), (y, -2, 2), opacity = 0.6, aspect_ratio = [1, 1, 10])

    ------------------------------------------------------

                                                      

    함수의 극댓값, 극솟값을 구하는 문제가 바로 컴퓨터(GDM)가 빠르게 잘하는 것이고, 이것이 바로 우리가 인공지능을 활용해서 합리적인 의사결정을 할 때, 우리에게 optimal solution을 주는 결정적인 알고리즘이 됩니다.


    The problem of finding a local maximum and local minimum value of the given function can be done well by the computer in a short time. This GDM algorithm becomes a critical tool to find an optimal solution that helps us to make a reasonable decision in machine learning.


      위의 그래프에서 산의 정상에 해당하는 부분에서 2 변수 함수는 극댓값을 갖게 되고, 계곡의 바닥에 해당하는 부분에서 극솟값을 갖게 됨을 직관적으로 확인할 수 있습니다, 이 점을 찾기 위해서는 일변수 함수의 도함수에 해당하는 개념이 필요한데, 이를 다변수 함수의 그래디언트(gradient)라고 합니다.


    The graph of in the figure above intuitively shows that has a local maximum value at the peak and a local minimum value at the valley. To find those critical points, the notion of a 'gradient of a multi-variable function' is required.


      2  변수 함수 의 그래디언트는 다음과 같이 정의합니다.


    The gradient of a function of two variables is defined as following.

                  grad


     이것은 2×1 벡터인데, 여기서 는 를 변수 에 관하여 편미분한다는 뜻으로 를 제외한 다른 변수는 모두 상수로 취급하여 미분하는 것과 같습니다. 도 마찬가지로 이해할 수 있습니다. 예를 들어, 의 그래디언트를 구하면 다음과 같습니다. (화면의 코드를 보세요)


    The gradient of is a 2×1 vector, where means the partial derivative of with respect to , that is, other variable except is considered as a constant. Similarly, means the partial derivative of with respect to y.


    [Example] The gradient of can be found as follows.


    ------------------------------------------------------------------

    var('x, y')  # variable

    f(x, y) = -x*y*exp(-x^2 - y^2)

    f(x, y).gradient()  # gradient

    ------------------------------------------------------------------

    (2*x^2*y*e^(-x^2 - y^2) - y*e^(-x^2 - y^2),

    2*x*y^2*e^(-x^2 - y^2) - x*e^(-x^2 - y^2))


    Answer : grad ((2*x^2*y*e^(-x^2 - y^2) - y*e^(-x^2 - y^2), 2*x*y^2*e^(-x^2 - y^2) - x*e^(-x^2 - y^2)) )


     i.e.


    grad ■

                 


    이것이 다변수함수의 그래디언트입니다.



    Gradient of multi-variable function

                              다변수함수의 도함수(그래디언트)


    마찬가지로 독립변수가 3개 이상인 다변수 함수에 대해서도 의 그래디언트를 구할 수 있습니다. 예를 들어, 변수 함수 에 대하여 그래디언트는 다음과 같습니다. (화면의 수식을 보세요)


    Similarly, the gradient of , a multi-variable functions with three or more variables, can be obtained. In general, the gradient of the n-variable function is as follows.

         grad

                                                         <---화면에 있으니 자막에는 없어도 됨)


    여러분은 앞으로 다변수 함수의 그래디언트를 바로 구해서 극값을 구하는데 활용할 수가 있는 것입니다.


    If you can find the gradient of a given multi-variable function, then you can use it to get the extreme values.


    최소제곱문제에서 첫 번째 예로 들었던 [5.2절]의 문제는 다음과 같은데, 각 데이터 에 대하여 를 일차함수 에 대입하여 얻은 값을 라 하고, 그러면  가 되겠죠. 이 선형연립방정식의 해가 존재하지 않는 경우에는. 각 데이터와 근사식 사이의 오차의 제곱 가 최소가 되는 , 를 구할 수가 있으며 데이터에 대하여 오차를 더한 오차함수는 다음과 같이 정의됩니다. (이때, 을 곱한 이유는 앞으로 2차식의 도함수를 계산하며 상쇄되도록, 단순히 계산의 편리성을 주기 위함이므로 크게 의미를 갖지 않으셔도 됩니다.) (화면의 수식을 보세요)


    In [Section 5.2], the first example of the least squares problem was this one. For each data , let's say is the value obtained by substituting in a linear function . So we have, . If the solution of this equation does not exist, it is possible to find , where the square of error is minimized. The error function , which is the sum of squared errors, is defined as follows: (The reason that we multiplied it by at the front of the error expression is only for the convenience of computation, so it doesn't affect anything on the conclusion.)


                                                             <---화면에 있으니 자막에는 없어도 됨)


    위의 오차함수에 대한 을 구하기 위하여, 경사하강법을 적용해 봅니다. 이때 초기조건 , 허용오차 , 학습률 로 두면 다음을 얻습니다. (화면의 코드를 보세요)


    Let's use the GDM to get the approximate solution that minimizes . Set an initial iterate , tolerance , and initial learning rate .


    --------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    var('a, b') 

    E(a, b) = 1/2*((a - 1)^2 + (a + b - 3)^2 + (a + 2*b - 4)^2 + (a + 3*b - 4)^2)

    gradE = E.gradient()

    u = vector([-2.5, -2.5])  # initial value

    tol = 1e-6 # tolorence 10^(-6)

    eta = 0.1  # learning rate

    r = []     # graph


    for k in range(300):

        g = gradE(u[0], u[1])

        gn = g.norm()

        r.append((u[0], u[1], E(u[0], u[1])))

        if gn <= tol:

            print("Algorithm success !")

            break

        u = u - eta*g


    print("u* =", u)

    print("E(x*) =", E(u[0], u[1]))

    print("iteration number =", k + 1)

    p1 = plot3d(E(a, b), (a, -3, 5), (b, -3, 5), opacity = 0.6)  #  E(a, b)

    p2 = line3d(r, color = 'red') + point3d(r, color = 'black', size = 50) 

    show(p1 + p2, aspect_ratio = [5, 5, 1]) # graph

    ---------------------------------------------------------------------

                                                             <---화면에 있으니 자막에는 없어도 됨)


    Algorithm success !

    u* = (1.49999929120196, 1.00000033198324)

    E(x*) = 0.500000000000364

    iteration number = 118   # number of iterations


    따라서 최소제곱직선은 이다.                           ■


    The solution is the same as that obtained from QR decomposition. The least square line is .


    앞에서 우리가 구한 것과 마찬가지로 주어진 데이터들로부터 알고리즘으로 <least square line>을 찾은 것입니다. 우리가 풀려던 least square problem의 해인 최소제곱 solution을 알고리즘으로 구해서, 그 답을 계수로 하는 최소제곱직선 을 얻은 것입니다.




    마찬가지로 [5.4절]에 있는 와 의 관계를 가장 잘 보여주는 (best fit) 이차함수 도 찾을 수 있습니다. 이때, 오차는 다음과 같습니다. (화면의 수식을 보세요)


    We can also find a quadratic function that best fits the given data. In this case, the error function is as following.


           

                  

                                                             <---화면에 있으니 자막에는 없어도 됨)


    역시 경사하강법을 사용하면 아래의 (best fit) 이차함수 을 얻습니다. 이때 초기조건을 , 허용오차 , 학습률 로 두면 다음과 같은 알고리즘을 얻게 됩니다. (화면의 코드를 보세요)


    We can use the same GDM code to find a best fit to below quadratic function . Take an initial value , tolerance and learning rate .


    ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    var('a, b, c')

    E(a, b, c) = 1/2*((a - 1)^2 + (a + b + c - 3)^2 + (a + 2*b + 4*c - 4)^2 + (a + 3*b + 9*c - 4)^2)

    gradE = E.gradient()


    u = vector([1.0, 1.0, 1.0]) 

    tol = 1e-6  # 10^(-6)

    eta = 0.01  # learning rate


    for k in range(5000):

        

        g = gradE(u[0], u[1], u[2])

        gn = g.norm()

                  

        if gn <= tol:

            

            print("Algorithm success!")

            break

        

        u = u - eta*g


    print("u* =", u)

    print("E(x*) =", E(u[0], u[1], u[2]))

    print("iteration number =", k + 1)

    ---------------------------------------------------------------------

                                                             <---화면에 있으니 자막에는 없어도 됨)


    Algorithm success!

    u* = (1.00000138155249, 2.49999723768073, -0.499999180018564)

    E(x*) = 1.59664737265671e-12

    iteration number = 4207


    즉,  (best fit) 이차함수 의 계수가 이므로, 최소제곱 곡선은 이다.                            ■

    Hence, the least square quadratic curve is . ■


     4207회를 거쳐서 이라는 최소제곱곡선을 구했습니다. 그래서 4개의 점을 지나는 best fit least square 곡선을 구한 셈이 되었습니다.


    After 4207 iterations, we now have as the least square curve. This means that we can always have a best fit least square quadratic curve that passes through four or more points.


    경사하강법 실습실 http://matrix.skku.ac.kr/math4ai-intro/W9/  에서 실습을 합니다.


    To practice the GDM algorithm, we can use following link.

     http://matrix.skku.ac.kr/math4ai-intro/W9/ 


    [Review] 


    인공지능과 최적해, 경사하강법과 최소제곱문제의 해 등 우리가 이번 주 배운 걸 복습해보면,


    경사하강법 알고리즘으로 극솟값을 구해보았습니다. 또, 응용문제, 최소제곱문제, curve fitting 하는 문제에 대해서 학습했습니다.


    In today’s lecture, we have learned how to find a local minimum using the GDM algorithm. We also learned some applications of it, which is a problem of finding the least squares solutions on curve fitting.


    함수의 그래프가 있을 때, 이 꼭짓점을 구하는 문제들을 학습했고 그것을 위해서 그래디언트가 필요하다는 것을 학습했습니다.


    We've seen the problems of finding the critical points of a multi-variable function. It could be done by using the gradient.


    그래서 f(x, y)의 그래디언트를 구하고, 그래디언트를 이용해서 주어진 함수의  error 함수를 정의하여, 경사하강법 알고리즘을 이용하여 최소제곱곡선을 구했습니다. 4207회 알고리즘을 거친 후 우리가 원하던 curve 을 최소제곱곡선으로 얻었습니다. 이것이 최소제곱곡선의 그래프입니다.


    For this least squares problem, we found the gradient of f(x,y), defined the error function , and then the least squares curve could be obtained by the GDM algorithm. After 4207 iterations, the least squares curve was obtained.


    오늘은 이렇게 실질적으로 우리가 언제든지 사용할 수 있는 경사하강법 알고리즘을 이론뿐만 아니라 코딩을 실습해서 실제 문제를 푸는 과정까지 학습했습니다.


    We learned not only the theory of gradient descent algorithm but also the practical coding to solve some real problems.


    오늘 배운 내용을 복습하고 마무리하도록 하지요.

    오늘 배운 경사하강법, 최소제곱문제의 해는 그야말로, optimization, 최적해를 구하는 가장 중요한 이론이기도 하고, 인공지능이 합리적인 의사결정을 하기 위해서 꼭 필요한 계산을 대신 해주는 알고리즘입니다. 이를 이해하기 위해서 9주차에는 경사하강법과 경사하강법의 응용으로 최소제곱 line, curve 구하는 과정과 극댓값, 극솟값을 배우는 이론, 알고리즘, 코딩을 실습했습니다.


    The GDM algorithm is the most important tool for finding optimal solutions. This is an algorithm that machine learning needs to make reasonable decisions. In this lecture, we learned the GDM and how to find the least squares line and curve by using GDM.


    경사하강법 알고리즘은 기본적으로 단계를 활용했습니다. 여기서 가 학습률이며 이것을 보통 0.01 정도 놓고 시작했습니다. 그래서 최솟값(minimum)을 구하는 문제를 학습했습니다. 기울기의 변화에 따라서 단계를 거쳐 극값에 수렴하는 과정을 학습했습니다.


    The iteration form of the gradient descent algorithm is . Here, is the learning rate, we usually start with 0.01. Through this process, the algorithm converges to an extreme value.


    여러분들은 학습하면서 http://matrix.skku.ac.kr/2020-Math4AI-PBL/ 에 있는 다른 학생들이 질문하고 답변한 내용들을 참고하고, 더 궁금한 내용은 문의게시판에 질문하시면 됩니다.


    For more information and discussion, please refer to the following link:

     http://matrix.skku.ac.kr/2020-Math4AI-PBL/


    여기서 미적분학을 마치고, 다음 시간에는 <인공지능에 사용되는 확률, 통계> 관련 내용을 학습하도록 하겠습니다. 수고 많이 하셨습니다.


    We're going to finish up this calculus chapter, and we will start  <Probability, Statistics background for AI> from the next week. Thank you.


           그림입니다.
원본 그림의 이름: K-001.jpg
원본 그림의 크기: 가로 560pixel, 세로 548pixel     그림입니다.
원본 그림의 이름: KakaoTalk_20200729_161558861.jpg
원본 그림의 크기: 가로 4032pixel, 세로 3024pixel
사진 찍은 날짜: 2020년 07월 29일, 오후 3:59   


     

    Week 10. Introductory Mathematics for Artificial Intelligence

     

                인공지능을 위한 기초수학 입문


    IV.  인공지능과 통계 (10-11주차)     108

    10. 순열, 조합, 확률, 확률변수, 확률분포, 베이지안(Bayesian)  108

    10.1 순열과 조합

    10.2 확률

    10.3 조건부확률

    10.4 베이즈 정리

    10.5 확률변수

    10.6 이산확률분포

    10.7 연속확률분포


    [10-1 pre]


    안녕하십니까. Math4AI 10주차 ‘인공지능과 통계’의 시작입니다.


    Welcome! In this week, we start to study Statistics.


    이제 인공지능에 필요한 ‘통계학의 기본적인 내용’을 학습하겠습니다.


    We will review the fundamentals of statistics needed for AI study.


    인공지능을 이해하기 위해서는 ‘선형대수학에서의 행렬’, ‘행렬분해’과, ‘미적분학에서의 도함수 및 그래디언트와 경사하강법’ 과 함께 ‘통계학에서의 확률지식과 확률변수, 공분산(covariance) 행렬’ 의 지식이 필요합니다. 이제 인공지능에 필요한 기초통계 지식에 대해서 2주에 걸쳐서 학습하도록 학습하겠습니다.


    In order to understand AI, it is required to have some knowledge of the following.

    - Matrix decompositions in linear algebra,

    - Function derivatives and the gradient descent algorithm in calculus, and probability,

    - Random variables, and covariance matrices in statistics.


    Henceforth, we will proceed with an introduction to the relevant statistical concepts needed for the study of AI. This expected to be covered within the next two weeks.

                   http://matrix.skku.ac.kr/math4ai-intro/W10/


    먼저 경우의 수를 세는 순열과 조합에 대하여 학습하고, 이를 바탕으로 확률에 관하여 정의합니다.


    Firstly, we study the concepts of 'permutations and combinations' which focuses on counting the number of possible outcomes in various situations or context. This leads to the definition of probability.


    확률은 전체 사건이 일어날 수 있는 경우의 수 분에 특정 사건이 일어날 수 있는 경우의 수로 이해할 수 있으므로, 경우의 수를 세는 것(counting technique)이 기본입니다.


    Probability can be understood as the ratio of the number of cases in which a particular event can occur and the number of cases in which the entire event can occur. Therefore various counting techniques become important.


    이어서 데이터 분석에서 중요한 개념인 ‘조건부 확률과 베이즈 정리’에 대하여 학습합니다. 베이즈 정리는 주어진 조건에서 어떤 현상이 실제로 나타날 확률을 구하는 방법으로, 불확실성 하에서 의사결정 문제를 수학적으로 다룰 때 중요하게 이용됩니다. 그리고 마지막으로 확률변수와 확률분포의 개념에 대하여 학습합니다.


    Next, we will study the concepts of <Conditional Probability and Bayes' Theorem>, which are one of the most important rules of probability theory applied to data analysis. Bayesian theorem is a method of finding the probability that an event/phenomenon will actually appear under a given condition, and is important when we analyze a decision making problem under uncertainty.


    Finally, we will study the concepts of random variables and probability distributions.


    순서는 1절에서는 순열과 조합, 2절에서는 확률, 3절에서는 조건부확률, 4절에서는 베이즈정리, 5절에서는 확률변수, 6절에서는 이산확률분포, 7절에서는 연속확률분포에 대해서 학습하겠습니다.


    We will study :

    Section 1, permutations and combinations

    Section 2, probability

    Section 3, conditional probability

    Section 4, Bayes' theorem

    Section 5, random variables

    Section 6, discrete probability distribution

    Section 7, continuous probability distribution..


    먼저 순열, 조합은 여러분들이 이미 익숙한 내용일 거고. 확률은 전체 사건 n(S) 분에 특정 사건 A가 일어날 경우의 수 n(A)가 바로 p(A)가 됩니다. 이를 계산하기 위해서 전체 사건 S 중에서 B의 사건이 일어날 경우의 수는 그림에서 보듯이 와 의 교집합(intersection) 5개를 각각 구해서 모두 더해주는 식으로 전체 사건이 일어날 경우의 수를 구할 수가 있습니다.


    First of all, we assume that everyone is familiar with the basic concepts of permutation and combination.The probability p(A) is the ratio of the number n(A), possibility that a specific event A can occur, and the number n(S), possibility that all events can occur.


     그리고 조건부 확률 , 즉 가 일어난다는 가정 하에 도 일어날 경우의 수를 구하는 식

                 그림입니다.


    을 이용해서 그 식을 일반화시키면 , 즉 가 일어난다는 가정 하에서 이 일어날 확률, 또 가 일어난다는 가정 하에서 가 일어날 확률 등을 다음과 같은 식

            그림입니다.

          그림입니다.

          그림입니다.

    으로 계산을 쉽게 할 수 있습니다. 여기에 베이즈 정리가 활용되는 것입니다.


    그림입니다. 그림입니다.



    Conditional probability is the probability of one event occurring under the assumption one or more other events have already occurred. Bayes' theorem is a formula that describes how to update the probabilities under the delicate hypotheses. It follows simply from the axioms of conditional probability. This formula make the computation easy and clear.


    이어서 확률변수, 확률분포에 대해서 학습을 하겠습니다.

    In the 2nd lecture, we will study <random variables and probability distribution>.


    [10주차 1강] http://matrix.skku.ac.kr/math4ai-intro/W10/


    초벌 번역; https://www.google.com/search?q=%EB%B2%88%EC%97%AD%EA%B8%B0


    여러분 반갑습니다. 이제 [인공지능을 위한 기초수학 입문]에서 중요한 부분인 행렬 부분, 도함수 부분을 마치고  <인공지능과 통계>, Part 4에 왔습니다. 이 파트4에서는 <순열과 조합, 확률, 확률변수, 확률분포, 베이지안> 등에 대해서 학습합니다.


    Welcome!! From this week, we start part 4, <Artificial Intelligence and Statistics>. In this lecture, we will study on <Permutation, Combination, Probability, Random Variables, Probability Distribution, and Bayesian>.


    먼저 이번 10주차 1차시에는 1절 <순열과 조합>으로 시작하죠.


    In the first session, we start with <Permutation and Combination>.


    경우를 세는 방법은 크게 두 가지, <순열과 조합>이 대표적입니다. 먼저 순열(permutation)은 순서를 고려하여 나열하는 경우의 수를 의미합니다. 예를 들어, 1, 2, 3, 4, 5가 적힌 5장의 카드 중에서 세 장을 택하여 순서대로 나열하는 경우의 수를 생각해보죠.

     

    There are two counting methods, <Permutation and Combination>. An arrangement in order is called a permutation. For example, consider the number of choices in which three cards are selected and listed in order from the five cards of 1, 2, 3, 4, and 5.


          (5장의 카드 중 하나) × (남은 4장의 카드 중 하나) × (남은 3장의 카드 중 하나)

    (One of five cards) × (one of four remaining cards) × (one of three remaining cards)


    전체 경우의 수는 첫 번째는 다섯 가지 가능성이 있었고, 두 번째는 네 가지 가능성이 있었고 마지막은 세 가지 가능성이 있었으니까 , 60가지의 경우가 있을 수가 있습니다. 3장을 택해서 순서대로 나열하는 경우입니다.


    In this case, we have to select 3 cards and arrange them in order. So for the total number of choices, the first position can have 5 choices, the second position have 4 choices, and the last position can have 3 choices, so we can have 60 different choices ().


    그 이유를 설명해보겠습니다.

    Let me explain this.


    카드 세 장을 순서대로 나열한다고 할 때 ⓵번 카드에 올 수 있는 숫자는 1, 2, 3, 4, 5로 5개가 있고, ⓶번 카드에 올 수 있는 숫자는 ⓵번 카드에 사용된 숫자를 제외하고 4가지이며, ⓷번 카드에 올 수 있는 숫자는 ⓵번과 ⓶번 카드에 사용된 2개의 숫자를 제외한 세 가지가 있습니다. 이들을 모두 곱하면 가지나 됩니다.


    Assuming that three cards are arranged in order, the number that can be appeared on the position ⓵ is 1, 2, 3, 4, and 5.


     Now only 4 numbers (except one card which is already at the position ⓵) are left be chosen for the position ⓶. Then only 3 numbers (except two card which are already at the positions ⓵ and ⓶) are left to be chosen for the position ⓷.


    이와 같이 서로 다른 개에서 개를 택하여 순서대로 나열한 순열의 수를 로 쓰고 다음과 같이 계산합니다.

     

    In this way, the number of ways of selecting and arranging k objects from among n distinct objects is


                       ()


    특히 일 경우에는, , n factorial 이 됩니다. 이걸 우리는 의 계승, factorial이라고 합니다. 를 계승을 이용하여 표현하면

    간단히 로 씁니다.


    Especially in the case of , = n!. We call this as n factorial.   can be simply expressed as using factorial.


    [예제 1]  이제 1부터 9까지의 숫자 중에서 서로 다른 3개를 선택하여 3자리 수를 만들려고 할 때, 만들 수 있는 자연수의 개수를 한번 구해볼까요.


    [Example 1] When we make a 3-digit number by selecting 3 different numbers from 1 to 9, what is the number of possible natural numbers that can be made?


    그러면 처음에는 9가지를 고를 수 있고, 그 다음 8가지, 그 남은 자리는 7가지 경우가 있으니까 9, 8, 7 즉, 9개 숫자에서 3개를 순서를 줘서 뽑는 것입니다.

    따라서 가 되겠습니다.

    손으로 계산하는 대신에 명령어를 이용하면


    We have 9 choices at the first position. Then 8 and 7 choices for the remaining digits, respectively, so 9x8x7 will be the answer.

    So it will be .


    Instead of computing it by hand, we use the 'Code' as follows:


     ------------------------------------

    factorial(9)/factorial(6) 

    ---------------------------------------------------------------------

    답: 504       # 이다.                     ■


    즉, 하면 504가 되는 것을 바로 확인할 수 있습니다. 복잡한 문제는 바로 명령어를 사용하시면 됩니다.


    In other words, we can immediately see that the answer is 504. If necessary, please do not hesitate to use the ‘Code' in http://matrix.skku.ac.kr/KOFAC/ or http://matrix.skku.ac.kr/math4ai-intro/W10/.



    이제 combination, 조합입니다. 조합은 순서와 상관없이 선택하는 경우의 수를 말하는데, 예를 들어서, 1, 2, 3, 4, 5가 적힌 5장의 카드에서 세 장을 택하는 경우의 수는 구하면 다음과 같습니다.


    A combination is a selection of all or part of a set of objects, without regard to the order in which they were selected. For example, if we choose  3 cards from 5 that the numbers 1, 2, 3, 4, 5 are written on it, the answer will be .


    즉, 5×4×3을 고른 후에, 그 3가지 장의 순서를 무시해도 되니까 3!로 나눠주면 됩니다.


    Because, after choosing 5×4×3, we just ignore the order of the 3 cards, by dividing 5×4×3 by 3!


    이렇게 구하는 이유를 다음과 같이 설명할 수 있습니다. 앞서 순열에서는 세 장을 택하는 경우마다 그 순서를 달리하면 모두 다른 경우로 여겨진 건데, 조합에서는 이미 택한 세 장에 대하여 순서대로 나열하는 것은 모두 같은 경우로 판단하므로 순열의 수를 세 장을 배열하는 수로 나누어주면 되는 것입니다. 3!로 나누어주면 우리가 원하는 답을 얻게 됩니다.


    The reason for this can be explained as follows. As mentioned in the permutation case, the order of each of the three cards was given, so all were considered to be different cases. However, in this combination case, the order of three cards need not to be considered. So we just divide by k!. That will give us the answer we want.


    이와 같이 서로 다른 개에서 개를 택하는 조합 combination의 수를 n choose k, 로 나타내고, 다음 공식에 의하여 계산합니다.

     

    In this way, the number of combinations for choosing k objects from n  different objects is expressed as n choose k, , and is computed by the following formula.    ()


    는 로 쓰기도 합니다.


       The  can be also written as .


    [예제 2]에서 500개의 넥타이로부터 5개의 넥타이를 택하는 방법의 개수는 입니다. 그런데 이 숫자가 꽤 크니까, 손으로 구하기 힘들 겁니다. 그러면 가르쳐준 명령어를 그대로 쓰셔서 이 주소에서 실습하시면 됩니다.


    [Example 2] The number of ways to select 5 ties from 500 ties is                            since .


     But since this number is quite large, it will be wise to use a 'text Code' as follows.


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    factorial(500)/(factorial(5)*factorial(495))

    --------------------------------------------------------

    255,244,687,600    #  (2천552억...)                                   ■


    하면 다음과 같이, 255,244,687,600나오는데, 이 숫자가 어떤 정도냐면, 2천552억4천4백6십8만7천6백, 엄청나게 큰 숫자입니다.


    이 명령어를 factorial을 쓰지 않고 binomial coefficient 이항계수를 구하라고 간단하게 binomial 명령어를 사용해도 됩니다.


    If we compute , it comes out 255,244,687,600. this number is over  255 billion, which is a huge number.


    We can do it in a simpler way by using a different code. The following binomial command gives the binomial coefficient that we want without using the above factorial command.


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    binomial(500, 5)  # binomial(n, k)  서로 다른 n개에서 k개를 택하는 조합의 수

    ---------------------------------------------------------------------

    255,244,687,600    #  (2천552억...)                                   ■


    마찬가지로 2천552억4천4백6십8만7천6백이라는 답을 얻을 수가 있습니다.


    Likewise, we will get the same answer of 255,244,687,600.


    지금까지는 선택할 때 중복을 허락하지 않은 경우를 소개하였습니다. 만일 중복을 허락한다면 다음과 같이 중복순열과 중복조합을 계산할 수 있습니다.

    여러분들에게 왜, 경우의 수를 세는 방법을 설명해주냐 하면, Counting Techniques 이라고 하는데요. 확률을 계산하려면 전체 사건의 경우 수와 특정사건의 경우 수를 세어야 하기 때문에 Counting Techniques 인 순열, 조합, 중복순열, 중복조합이 필수적인 도구가 됩니다.


    So far, we have introduced the cases where repetition is not allowed when we select. If we allow repetition, we can calculate <Permutation with repetition> and <Combinations with repetition> as follows.


    Now I explain how to count the number of cases, which is called <the Counting Techniques>. In order to compute the probability in AI, we have to count the total number of events and the number of specific events. So Counting Techniques such as permutation, combination, <Permutation with repetition>, and <Combinations with repetition> are essential tools in AI.


    중복순열(repeated permutation), 를 설명하겠습니다. 중복순열은 서로 다른 개에서 중복을 허락하여 개를 택하여 순서대로 나열한 경우이고 다음 공식에 의하여 계산합니다.


    Let me explain the <Permutation with repetition>.  It is a case when k objects are selected from n different objects and listed in order by allowing <repetition>, and is calculated by the following formula, .


    설명하면, n개 중에서 k개를 택하여 순서대로 나열하려면 첫 번째 올 수 있는 경우의 수가 n이고, 두 번째 올 수 있는 경우도 n이고, 세 번째 올 수 있는 경우의 수도 n입니다. 그러면, n을 k번 곱해준 것과 똑같은 경우가 됩니다. 굉장히 간단하죠. 중복순열은. 제일 간단하네요.


    It's very simple. When we select k out of n objects and list them in order, there are n choices for the first position and there are n choices for the 2nd position and there are n choices for the 3rd position etc. This means the answer is n^k. The <Permutation with repetition> is the simplest one.


    그 다음이 중복조합(repeated combination)은 라고 씁니다. 서로 다른 개에서 중복을 허용하여 순서 없이 개를 택하는 경우의 수는 아래 공식인, 입니다. 


    If we allow repetition, the number of choices when we select k out of n different objects and list them without order is .

    [This is a basic question of how to distribute ݑ› objects (call them "balls") into ݑ˜ categories.]


    [Suppose we have ݑ›+ݑ˜−1 distinct balls and two bags. Now we want to .\PICk ݑ˜ balls and put them in bag 1 and the rest ݑ›−1 balls in bag 2.


    If we choose ݑ˜ balls first, there are possible ways to do that. On the other hand, we can think you are actually .\PICking ݑ›−1 balls first and then put them in bag 2 and the rest ݑ˜ balls in bag 1. These two methods should be equivalent (choosing ݑ˜ balls for bag 1 is the same as choosing ݑ›−1 other balls for bag 2). Therefore, = . ]


    [예제3]을 풀어보면, 숫자 1, 2, 3, 4, 5 중에서 중복을 허락하여 세 개를 택해 일렬로 나열하여 만든 세 자리의 자연수가 5의 배수인 경우를 구하여라. 그러면, 먼저 세 자리 자연수의 일의 자리를 5로 고정 시킵니다. 빈 두 자리는 1, 2, 3, 4, 5 중에서 중복을 허락하여 나열하는 경우의 수와 똑같으니까, 서로 다른 n=5에서 중복을 허용하여 k=2개를 선택하는 중복순열의 수를 구합니다. 인 것을 쉽게 확인할 수 있습니다.


    [Example 3] Find the number of cases that the natural number of three digits is made of allowing repetitions from the numbers 1, 2, 3, 4, 5, is listed in a row, and is a multiple of 5.


    Answer: First, fix the last digit by 5, (*)(*)(5). Then it is a problem of '2 combinations out of 5 with repetitions' which is = 25. We can easily find it.


    [예제4] 4명의 사람이 A, B, C 중 한 명에게 무기명으로 투표를 할 때, 나올 수 있는 경우의 수가 몇 가지가 있을까요?

    4명이 무기명으로 투표하는 방법은 4명이 다 A를 찍을 수도 있고, 4명이 다 C를 찍을 수도 있고, 섞어서 찍을 수도 있죠. 이 경우는 서로 다른 3개에서 중복을 허용하여 4개를 선택하는 중복조합의 수와 같습니다. 그래서, 15가지인 것을 구할 수 있습니다.


    [Example 4]  Find the number of cases when four people vote anonymously for one of A, B, or C.


    Answer. In any anonymous voting, all 4 people can vote for A, and all 4 people can vote for C, etc. So this is the case of <3 combinations of 4 items with repetition>. Now we have `  _{n} `H`  _{k} `= ` `  _{3} `H`  _{4} `= ` _{3+4-1} C _{4}`=`  _{6} C  _{4} =15 . The answer is 15 cases.


    이렇게 4가지 기본적인 Counting Technique을 배운 후 이걸 이용해서 확률을 학습합니다.


    After we cover this 4 basic Counting Techniques, we are ready to learn probability.


      Section 10.2 Probability.


    특정 사건(event)이 일어날 가능성을 수 0과 1 사이의 값으로 나타낸 것을 확률( probability)라고 합니다. 예를 들어서, 동전 던지기를 한 번 했을 때 앞면이 나올 확률은 이고, 확률이 0이라는 의미는 사건이 절대로 일어날 수 없음을 의미하며, 확률이 1이라는 것은 그 사건이 반드시 일어난다는 것을 의미합니다. 


    The probability of an event is a number between 0 and 1, where, roughly speaking, 0 indicates impossibility of the event and 1 indicates certainty. For example, if we flip a coin once, the probability of getting a head is , the probability of 0 means that the event will never happen, and the probability of 1 means that the event must happen.


    사건이 일어날 확률을 수학적으로 분석하기 위해서는, 먼저 어떠한 사건들이 발생 가능한지를 명확히 알아야 합니다. 예를 들어서, 동전 던지기의 경우에 발생 가능한 사건들은 앞면이 나오는 경우, 뒷면이 나오는 경우 두 가지밖에 없죠. 또, 정육면체 주사위의 경우는 1, 2, 3, 4, 5, 6으로 나올 수 있습니다. 이러한 사건들의 집합을 표본공간, 표본공간(Sample Space)이라고 합니다. 이와 같이 동전 던지기의 sample space는 {앞면, 뒷면}이고, 정육면체의 주사위의 sample space는 {1, 2, 3, 4, 5, 6}입니다.


    To analyze the probability of an event mathematically, we should know clearly what events are possible. For example, in the case of a coin flip, there are only two possible events: heads and tails. Also, in the case of rolling a die, there are 6 possible events: 1, 2, 3, 4, 5, 6. This set of all possible events is called a sample space. Likewise, the sample space of a coin flip is {heads, tail}, and the sample space of rolling a die is {1, 2, 3, 4, 5, 6}.


    이제 확률을 정의해보죠. 어떤 실험이나 관찰에서 각 경우가 일어날 가능성이 같을 때, 일어나는 모든 경우의 수를 n(S)로 쓰고, 어떤 특정한 사건 A가 일어날 경우의 수를 n(A)라고 하면, 사건 A가 일어날 P(A)는 다음과 같습니다.

     


    Now let's define the probability. In some experiments or observations, if we denote the total number of possible outcomes n(S) and denote the number of ways event A can occur by n(A), then P(A) is defined by

     

                         



    이것을 기하학적으로 이해하면 인 에 속할 확률은


    가 됩니다. 앞에서는 counting을 해서 확률을 구한 것이고, 기하학적으로는 영역의 면적을 구해서 확률을 구한 것으로 이해할 수 있습니다.



    We can geometrically understand this concept with the areas, the probability is  


    In this formula, the probability was obtained by counting the number of cases. It can be understood geometrically as the probability can be found by the area of ​a region.


    다음은 통계적 확률입니다. 어떤 시행을 번 반복하였을 때, 특정사건 가 일어난 횟수가 번 이라할 때, n이 한없이 커짐에 따라 상대도수 k/n가 일정한 값 p에 가까워지면 이 값 p를 사건 A의 통계적 확률이라고 합니다. 우리가 시행을 통해 얻은 확률은 시행횟수가 충분히 커지면 수학적으로 얻은 P(A)에 수렴하게 되어있습니다. . 이걸 대수의 법칙 이라고 합니다.


    The next to.\PIC is a statistical definition of probability. The P(A)=p is called the statistical probability of event A when the relative frequency k/n approaches a constant value p as the number of trials n approaches infinity. The probability we obtained through trials converges to the statistical probability P(A) obtained mathematically when the number of trials is sufficiently large. This is called the Law of large numbers.


    위에 정의한 확률들은 다음 성질들을 만족합니다.


    사건 의 확률을 라 하면

    ① 표본공간 에서 임의의 사건 에 대하여 이 성립한다.

    ② 표본공간 에 대하여 (표본공간 전체의 확률은 1)이 성립한다.

    ③ 공사건 에 대하여 이 성립한다.

    ④ 두 사건 , 가 동시에 발생하지 않는 배반사건이면 다음이 성립한다.

                     

    ⓹ 사건 가 일어나지 않는 경우를 이라 하면 이 성립한다.


    The probability satisfies the following properties:


    Let be the probability of an event A.

     

    ① for arbitrary event in the sample space.

    ② For , .

    ③ for the empty set.

    ④ If the two events A and B do not occur at the same time,

                     

    ⓹ If the is the event that does not occur, then .


    [예제 5] 주머니 속에 검은 공이 3개, 흰 공이 2개, 붉은 공이 1개 들어있을 때, 여기서 동시에 두 개의 공을 꺼낼 때, 같은 색의 공이 나올 확률은 무엇인가 하는 것입니다.


     먼저 전체 경우의 수를 구해 놓습니다. 6개의 공 중에서 두 개의 공을 꺼내는 경우의 수는이죠. 다음은 같은 색인 경우를 구합니다. 검은 공이 2개 나올 경우와 흰 공이 2개 나올 경우인데, 검은 공 3개 중에서 2개를 고르는 경우의 수와 흰 공 2개 중에서 2개를 고르는 수가 동시에 일어나지 않으니까, 위의 합집합(union)에 대한 성질을 사용하여 가 됩니다. 그래서 사건이 일어날 확률은 4/15가 되겠습니다. binomial 명령어를 사용하면 4/15가 바로 얻어집니다.


    [Example 5] Suppose that we have 3 black balls, 2 white balls, and 1 red ball in a pocket. What is the probability of choosing two balls of the same color when we take out two balls at the same time from the pocket?


    First, find the total number of cases. The number of cases that we take out two out of the six balls is . Then counting the number of cases with the same color. There are only two cases of choosing two black balls and two white balls. Choose 2 out of 3 and choose 2 out of 2. And add them to have 4 cases. So the probability  is 4/15. With the binomial 'Code', we have the answer 4/15 right away.


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    ( binomial(3, 2) + binomial(2, 2) ) / binomial(6, 2)

    ---------------------------------------------------------------------

    4/15                                                          ■


    [예제 6] 아래의 R 명령어로 동전 한 개를 10회 던져 보고, 뒷면의 수와 앞면의 수를 기록해 봅시다. 동일한 방법으로 같은 동전을 100회 던지는 실습을 한번 해보죠. 여기서는 R 명령어를 사용했는데요. R 명령어를 사용할 때는 이 웹 주소에 가서 이걸 copy해 놓고 우측 아래의 명령어를 파이썬이나 Sage 대신에 R로 지정하고 클릭하시면 됩니다. 그러면 앞, 앞, 앞, 뒤, 뒤, 뒤, 뒤, 뒤, 앞, 뒤가 나옵니다.


    [Example 6] Flip a coin 10 times with the R command below, and find the number of 'Tail' and the number of 'Head'. Let's practice to flip the same coin 100 times in a similar way. We may use <R code> here. If we want to use a <R code>, visit this link http://matrix.skku.ac.kr/math4ai-intro/W10/ and practice with 100 instead of 10. We choose R instead of Python or Sage in the lower right corner of the Sage cell, and just click ’Evaluate’ button.  Then we will see the output.


    ---------- http://matrix.skku.ac.kr/KOFAC/  ---  < R  명령어> ---------

    coin = sample(c("뒤","앞"), 10, replace = TRUE)   # 10회 반복 < R  명령어>

    coin

    ---------------------------------------------------------------------

    [1] "앞" "앞" "앞" "뒤" "뒤" "뒤" "뒤" "뒤" "앞" "뒤"                ■


    명령어를 coin으로 하지 않고 table(coin) 이라는 명령어를 주어서 표로 표시하면, 이런 식으로 printout이 됩니다.


    If we use a new code called ‘table(coin)’ in R environment (instead of the Code ‘coin’), then we will have a simple table that shows the number of events in each case.


    그림입니다.
원본 그림의 이름: CLP000034380009.bmp
원본 그림의 크기: 가로 472pixel, 세로 130pixel ---------- http://matrix.skku.ac.kr/KOFAC/  ---  < R  명령어> ---------

    coin = sample(c("뒷면","앞면"), 100, replace = TRUE)   # 100회 반복 < R  명령어>

    table(coin)

    ---------------------------------------------------------------------

    coin

    뒷면 앞면

    51   49                                                       ■


    100번을 반복하면 coin이 뒷면이 51번 나오고 앞면이 49번 나옵니다. 또 시행해보면 random하게 test해보는 것이기 때문에 약간씩 달라집니다.


    When we repeat 100 times, the 'Tail' came out 51 times and the 'Head' came out 49 times. If we try again, we will be slightly different output because it is a random generator.



    이게 10번, 100번, 1000번, 10000번 반복하면, 점점 뒷면이 나올 확률과 앞면이 나올 확률, 경우의 수들이 비슷하게 나오는 걸 확인할 수 있습니다. 이와 관련된 내용은 통계학 명령어들, R 명령어들을 모아놓은 실습실에서 실습하실 수 있습니다.

    http://matrix.skku.ac.kr/2018-album/R-Sage-Stat-Lab-2.html 


    If we repeat this same test, 10 times, 100 times, 1000 times, 10,000 times, we will see that the probability of events for 'Tail' and "Head' are getting similar to each other. We can practice this in a lab R commands.

    http://matrix.skku.ac.kr/math4ai-intro/W10/ 


    그래서 시행의 횟수를 아주 크게 늘려 가면서 앞의 대수의 법칙에서 설명한 바와 같이 뒷면과 앞면이 나오는 확률이 , 수학적 확률 로 수렴함을 확인할 수 있습니다. 이런 식으로 여러분들은 실습을 코드를 이용해서 다양하게 해보실 수가 있습니다.


    If the number of trials are sufficiently large, we will see that the probability of each event converges to 1/2, which was the mathematical probability as the Law of large numbers told. We can do a variety of practice with this simple Code.


    다음 [예제 7]은, 1000개의 제품 중에 불량품이 3개가 있는데, 이 제품 중에서 10개의 제품을 구입했을 때 다음 두 가지를 구하는 겁니다. 첫 번째는 구입제품 중 불량품이 한 개도 없는 경우고, 두 번째는 구입제품 중 불량품이 적어도 한 개 이상 있는 경우, 각각 한번 구해보죠.


    In [Example 7], there are 3 defective products out of 1000 products. When 10 of these products are purchased, Find the following two probabilities.

    (1) There is no defective product among the purchased products.

    (2) There is at least one defective product among the purchased products.


    구입 제품 중에 불량품이 한 개도 없는 경우는, 일단 1000개의 제품 중에 10개의 제품을 선택하는 경우의 수는 이죠. 많은 경우의 수가 있습니다. 첫 번째 불량품이 한 개도 없는 경우는 정상 제품인 997개에서 10개를 모두 선택하고, 불량품 3개에서는 하나도 선택하지 않는 경우 밖에 없으므로 그 경우의 수는 입니다.  Sage 코드를 이용하여 계산하면 다음과 같습니다. (화면의 코드를 보세요)


    (1) The number of cases where 10 products are selected out of 1000 products is . If there is no defective product, then all 10 items should be selected from 997 normal products, and none of the three defectives are selected. so the number of cases is . Calculation using Sage code is given as follows. (See the code on the screen)



     ---------- http://matrix.skku.ac.kr/KOFAC/

    # 모두 불량품이 아닐 확률

    (binomial(997, 10)*binomial(3, 0)/binomial(1000, 10)).n(digits = 7)

    ---------------------------------------------------------------------

    0.9702695                                            ■


    digits=7은 답이 소숫점 아래로 길게 내려갈 수 있으니까, 유효숫자가 7자리에서 멈춰달라는 의미입니다. 97% 정도가 나오게 되지요. R코드를 이용해서 계산하면 똑같은 내용인데, 이렇게 간단하게 쓸 수 있습니다.


    Here, the optional digits argument specifies the number of decimal digits. For example, ‘Digits=7’ in the code above specifies 7 decimal digits. We have the same answer if we compute by using R code as follows.


    ---------- http://matrix.skku.ac.kr/KOFAC/  ---  < R  명령어> ---------

    # 모두 불량품이 아닐 확률 < R  명령어>

    choose(997, 10)* choose(3, 0)/ choose(1000, 10)

    ---------------------------------------------------------------------

    [1] 0.9702695                                          ■


    997개 중에서 10개를 고르고, 3개 중에서 0개를 고르는 것을 1000개 중에서 10개 고른 걸로 나눠라. 바로 답이 97%로 나오죠. 쉽게 확인할 수 있습니다.


    choose(997, 10)* choose(3, 0)/ choose(1000, 10)

    gives the answer, which is 9702695% as we can easily check.


    그 다음에 불량품이 적어도 한 개 이상 있을 확률은 무엇일까요. 그러면 1에서 불량품이 하나도 안 나올 확률을 빼주면 되지 않겠어요. 그러니까 앞에서 구한 불량품이 하나도 안 나올 확률을 1에서 빼주면 적어도 불량품이 하나는 있는 경우가 되겠죠. 그것을 Sage 코드를 이용해서 확인해보면 이렇습니다.


    (2) For the case of that there will be at least one defective item, we just subtract the above probability (with no defective products) from 1. That will be the probability of having at least one defective product. If we check that with 'Sage Code', it is given as follows.


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    # 적어도 불량품이 1개 이상 있을 확률

    (1 - binomial(997, 10)*binomial(3, 0)/binomial(1000, 10)).n(digits = 7)

    ---------------------------------------------------------------------

    0.02973045                                                     ■


    1에서 불량품이 하나도 안 나올 확률을 빼주면 2.97% 정도가 되네요. 이것을 R 코드를 사용해서 계산하면 다음과 같이 2.97%. 이렇게 나오네요.


    It was about 2.97%. Compute this using R code is 2.97% as follows. It comes out like this.


    ---------- http://matrix.skku.ac.kr/KOFAC/  ---  < R  명령어> ---------

    # 적어도 불량품이 1개 이상 있을 확률 < R  명령어>

    1 - choose(997, 10)*choose(3, 0)/choose(1000, 10)

    ---------------------------------------------------------------------

    [1] 0.02973045                                                ■


    여러분들은 이와 같이 복잡한 계산을 코딩을 이용해서 컴퓨터 또는 인공지능에게 시켜서 계산을 해본 것입니다. 이상 확률에 대한 설명을 마치고, 조건부 확률을 공부한 뒤에 이번 시간을 마치겠습니다.


    We now have done some complicated computation by using a simple 'Code' on the web with your computer/AI. Now we finish up this probability part and start to explore the 'Conditional probability'.


    조건부 확률은 확률과 데이터 분석에서 사용되는 아주 중요한 개념인데, 어떤 사건 가 일어났다는 조건하에서 사건 가 일어날 확률입니다. 이것을 사건 에 대한 사건 의 조건부확률, conditional probability라 하고 로 표시합니다.


    The conditional probability is a very important concept in data analysis, which is the probability that an event B occurs under the condition that the event A occurred. This is called <the Conditional Probability> and is expressed as .


        (단, )

    이 경우에는 분모가 A의 확률이니까 로  0보다 커야 정의가 됩니다.


    In this case, the denominator, the probability of A, must be greater than 0.



    그림으로 그리면, P가 A가 일어나고 그 상황에서 B가 일어날 확률은 그 전체 확률 분에 이 작은 교집합(intersection)의 확률, 이 면적이 되는 겁니다. 그림이 쉽게 이해가 됩니다.


    This diagram shows that the probability is the probability of this small intersection over the total probability. The diagram helps us to understand this concept.


    그리고 조건부 확률의 정의로부터 다음의 곱셈정리 관계식을 얻을 수 있습니다.


    And from the definition of the conditional probability, we can get the following rule.


       


    일반적으로 사건이 여러 개가 있다면, 즉 n개의 사건이 있다고 그러면


    In general, we can generalize it for multiple events, such as n events.


    이렇게 쓸 수 있습니다. 그리고 도 일어나고 도 일어나고 도 일어나고 도 일어나는, 즉 모든 경우가 일어날 확률은 다음과 같이 쓸 수 있다는 의미입니다. 조건부 확률을 이용해서 이것을 구할 수 있다는 의미입니다. (화면의 수식을 보세요)


    This means that we can find easily. (See the formula on the screen)



      

        

              

                                                           <---화면에 있으니 자막에는 없어도 됨)


    [예제 8] 두 사건 , 에 대하여 , 일 때, 의 값을 구하기 위해 조건부 확률 식에 값을 대입하면 됩니다.


    [Example 8] For two events , with , , Find .


    Sol.) We can use the formula on the conditional probability.



    This is the answer, 16/21.


    는 드모르간의 법칙에 의해 이 됨을 바로 확인할 수 있습니다. 계산한 결과, 16/21로 바로 답을 구할 수 있습니다.


    We know that = by De Morgan's law. The answer is 16/21.


    이와 같이 확률에 대한 공식, 조건부 확률에 대한 공식을 이용해서 Counting 한 것들의 관계로부터 우리가 원하는 확률을 계산할 수 있습니다.


    10주차 2차시에는 조건부 확률에 대한 주요 정리인 베이즈 정리를 학습하도록 하겠습니다.

    수고 많이 하셨습니다.


    In this lecture, we can find the probability we want. In the second lecture of Week 10, we will study Bayesian theorem, the main theorem for conditional probability.  Good job^^. Thanks.


    [10-2 pre]


    반갑습니다. 지난 1차시에는 우리가 순열, 조합과 확률에 대해서 학습했습니다.


    Welcome! In the previous lecture, we learned about <permutations, combinations and probabilities>.


    이번 2차시에서는 이어서, ‘조건부 확률, 베이즈 정리, 확률변수, 확률분포’에 대해서 학습을 하고 실습하도록 하겠습니다.


    In this second lecture of Week 10, we will study ‘Conditional probability, Bayesian theorem, Random variables, and Probability distribution’.


    [10주차 2강] http://matrix.skku.ac.kr/math4ai-intro/W10/


    반갑습니다. 10주차 지난 시간에 <인공지능과 통계>에서 순열과 조합에 대해서 배웠습니다. 순열(permutation)을 구하는 방법을 배우고 504를 실제 구해봤습니다.

    또 조합으로 500개 넥타이에서 5개를 선택하는 방법을 배웠고, 이항계수(binomial coefficient)를 이용해서 계산하는 방법을 학습했습니다.

    수학적 확률, 기하적 확률, 통계적 확률과 대수의 법칙까지 학습을 했습니다. 그리고 확률의 성질들을 학습했습니다.

    counting하는 방법에 대한 code를 배웠고, 여기서 R 명령어로 랜덤(random)하게 동전을 고르는 방법을 배웠습니다. 예를 들어, 10개를 고르라고 명령하면 이렇게 10개를 고를 수 있습니다. (화면의 코드를 보세요)


    Welcome!!  In the last lecture, we have reviewed <Permutations>, <Combinations>, <Permutation with repetition> and <Combinations with repetition>, <Probability>, <Conditional probability>.


    In general, we studied about mathematical probability, geometric probability, statistical probability, properties of probabilities and the Law of Large number.


    In addition, we studied how to write 'Codes' to perform counting operations, and we learned how to randomly flip coins with 'R language'. For example, if we are asked to choose 10 samples, we can do it with the following Code.  (See the code on the screen)



     ---------- http://matrix.skku.ac.kr/KOFAC/  ---  < R  명령어> ---------

    coin = sample(c("뒤","앞"), 10, replace = TRUE)   # 10회 반복 < R  명령어>

    coin

    ---------------------------------------------------------------------

                                                             <---화면에 있으니 자막에는 없어도 됨)


    [1] "앞" "앞" "앞" "뒤" "뒤" "뒤" "뒤" "뒤" "앞" "뒤"                ■



    여기서, 숫자를 100으로 바꿔주면 100개가 골라집니다. 100개쯤 골랐더니 몇 개쯤 나오나 했더니, 여기서는 54개가 뒷면, 46개가 앞면쯤 나왔습니다. 다시 한 번 random 하게 실행을 해보면 50개, 50개, 정확하게 같은 값이 나오네요. 1000개, 10000개 수를 늘리면 점점 같은 개수로 수렴하게 될 겁니다. 그게 대수의 법칙(law of large number)입니다.


    Here, if we change the number of samples to 100, 100 samples will be randomly selected. When we toss a coin 100 times, there comes 54 tails and 46 heads at this trial. If we run it randomly again, we now have 50, 50. If we increase the number of trial to 1000 and 10000, it will gradually converge to the same number. That's the <Law of Large number>.


    [예제 7]에서 불량품을 선택하는 방법을 구해보았습니다. R 언어로 바꿔주고. 그 결과, 불량품이 하나도 나오지 않을 경우가 0.97, 97%였습니다. 그러면 불량품이 하나라도 있을 확률은 0.03, 3%쯤 되겠죠. 바로 계산이 됩니다.


    지난 시간에 여기까지 배웠고, 이번 시간에는 조건부 확률과 베이즈 등을 학습하도록 하겠습니다.


    In [Example 7], we tried to find out the probability of selecting non defective products. We computed this probability by using R-language, there were 97% probability of choosing none of the defective products. Then, the probability to have at least one defective product was about 3%. It can be computed easily after we understand problem clearly.


    조건부 확률은 확률과 데이터 분석에서 사용되는 아주 중요한 개념인데, 어떤 사건 가 일어났다는 조건하에서 사건 가 일어날 확률입니다. 이것을 사건 에 대한 사건 의 조건부확률(conditional probability)라 하고 로 표시합니다.


        (단, )

    이 경우에는 분모가 A의 확률이니까 겠죠. 0보다 커야지 정의가 되겠죠.


    The conditional probability is a very important concept in data analysis, which is the probability that an event B occurs under the condition that the event A occurred. This is called <the Conditional Probability> and is expressed as .


           (단, )

    In this case, the denominator, the probability of A, must be greater than 0.


    그림으로 그리면, P가 A가 일어나고 그 상황에서 B가 일어날 확률은 그 전체 확률 분에 이 작은 교집합(intersection)의 확률, 이 면적이 되는 겁니다. 그림이 쉽게 이해가 됩니다.


    This diagram shows the conditional probability is the probability of this small intersection over the total probability. The diagram helps us to understand the concept.


    그리고 조건부 확률의 정의로부터 다음의 곱셈정리 관계식을 얻을 수 있습니다.


    And from the definition of conditional probability, we can get the following formula.


       


    일반적으로 사건이 여러 개가 있다면, 즉 n개의 사건이 있다고 그러면


    In general, we can generalize it for multiple events, such as n events.


    이렇게 쓸 수 있습니다. 그리고 도 일어나고 도 일어나고 도 일어나고 도 일어나는, 즉 모든 경우가 일어날 확률은 다음과 같이 쓸 수 있다는 의미입니다. 조건부 확률을 이용해서 이것을 구할 수 있다는 의미입니다. (화면의 수식을 보세요)


    This means that we can find easily. (See the formula on the screen)



      

        

              

                                                             <---화면에 있으니 자막에는 없어도 됨)


    [예제 8] 두 사건 , 에 대하여 , 일 때, 의 값을 구하기 위해 조건부 확률 식에 값을 대입하면 됩니다. (화면의 수식을 보세요)


                                                             <---화면에 있으니 자막에는 없어도 됨)


    는 드모르간의 법칙에 의해 이 됨을 바로 확인할 수 있습니다. 계산한 결과, 16/21로 바로 답을 구할 수 있습니다.


    [Example 8] For two events , with , , find the probability .


    Sol.) We can use the formula on the conditional probability.



    This is the answer, 16/21.


    는 드모르간의 법칙에 의해 이 됨을 바로 확인할 수 있습니다. 계산한 결과, 16/21로 바로 답을 구할 수 있습니다.


    We know that = by De Morgan's law. The answer is 16/21.


    다음으로, 베이즈 정리(Bayes’ theorem)를 학습합니다.

    베이즈 정리는 주어진 조건에서 어떠한 현상이 실제로 나타날 확률을 구하는 방법으로, 불확실성 하에서 의사결정 문제를 수학적으로 다룰 때 아주 중요하게 사용되는 내용입니다. 특히, 정보와 같이 눈에 보이지 않는 무형자산이 지닌 가치를 계산할 때 유용하게 사용됩니다.

    이때, 사용되는 용어들을 먼저 정리하겠습니다. 사전확률(prior probability)은 관측자가 이미 알고 있는 사건으로부터 나온 확률을 말합니다. 그리고 사후확률(posteriori probability)은 사전확률과는 대비되는 개념으로 실제의 데이터나 조건이 부과되었을 때 기대되는 조건부 확률을 말합니다.

     

    Now we start the Bayes’ theorem.


    Bayes' theorem describes the probability of an event, based on prior knowledge of conditions that might be related to the event.

    It is very important when dealing with decision-making problems mathematically under uncertainty. In particular, it is useful when calculating the value of invisible and intangible assets such as information.

    First, we will start by introducing some frequently used terms. <A prior probability> of an event is the probability of the event computed before the collection of new data. And <A posterior probability> is the revised or updated probability of an event occurring after taking into consideration new information.


    즉, 조건부 확률은 어떤 특정 사건이 이미 발생하였는데, 이 특정 사건이 나온 이유가 무엇인지 불확실한 상황을 식으로 나타낸 것이며 로 표현할 수 있습니다. 여기서 는 이미 일어난 사건이고, 사건 를 관측한 후에 그 원인이 되는 사건 의 확률을 따졌다는 의미로 사후확률이라고 정의합니다. 베이즈 정리는 사전확률과 사건으로부터 얻은 자료를 사용하여 사후확률을 추출해 내는 것으로서 즉 사전확률과 사후확률의 관계를 조건부 확률을 이용하여 계산하는 이론입니다.


    다음 그림이 잘 설명하고 있는데요. 이 표본공간 를 분할(partition)을 했을 때, 임의의 사건 에 대하여 다음이 성립합니다.


    Conditional probability is used for explaining the situation in which a specific event occurred but no one knows why the event occurred. In , B represents an event that has already occurred, and is considered as the posterior probability, meaning that the probability of the event after the occurrence of B is observed. The Bayesian Theorem is a theory that extracts posterior probabilities using data obtained from prior probabilities and events, that is, it computes the relationship between prior probabilities and posterior probabilities using conditional probabilities.


    The following diagram explains it well. When this sample space is partitioned by , the following holds for any event .



    이때 는 서로 배반(exclusive)입니다. 따라서 각각의 확률을 더한 것과 같습니다. 그러면 확률의 곱셈정리로부터 아래 전확률 공식(Law of Total Probability)을 얻을 수 있는데, 그것은 P(B)가 다음과 같이 쓰여질 수 있다는 것입니다.



    이제, 임의의 에 대한 조건부확률을 가지고 우리가 앞에서 배운 이 내용을 이 공식에 대입하면 다음과 같은 식을 얻을 수 있습니다. 이게 베이즈 정리(Bayes’ theorem)입니다.  (화면의 수식을 보세요)


    Now, is mutually exclusive. So P(B) is equal to the sum of probabilities of each . Then from Multiplication theorem on probability, we can get <the law of total probability>, which means the following formula, P(B):



    Now, if we  use the definition of conditional probability and <the law of total probability>, we can get the following formula, which is called the Bayes' theorem. (See the formula on the screen)



     

             


                                                             <---화면에 있으니 자막에는 없어도 됨)


    베이즈 정리에서 를 사건 의 사전확률이라고 하고, 를 사건 의 사후확률이라고 합니다.


    베이즈 정리가 어디에 사용될까요?

    [예제 9]에서 보듯이 3대의 기계 가 각각 이 공장의 생산품 전체의 를 생산한다고 합시다. 그리고 이들 기계가 불량품을 생산할 비율은 각각 라고 합니다. 그러면 한 제품을 임의로 선택할 때 그 제품이 불량품일 확률을 구해봅시다. 또한 불량품이 기계 에 의하여 생산될 확률을 구하라는 문제를 생각해보겠습니다.

    풀이는 구입한 개의 제품이 기계 로부터 생산된 제품인 사건을 로 나타내고, 그것이 불량품이라는 사건을 로 나타내면, 제품을 생각하는 사건, 불량품을 생각하는 사건과 확률은 다음과 같습니다.


    In Bayes' theorem, is called the prior probability of an event and is called the posterior probability of an event .


    When is Bayes' theorem used?


    [Example 9]  Suppose that three machines produce the of products in this factory respectively. And the defective rate of A, B, C are , respectively.


    (1) Find the probability that the purchased product is defective when we randomly select one product.

    (2) Find the probability that this purchased defective product was produced by the machine .


    Solution)

    Suppose that an event in which the purchased product is produced by a machine is denoted by , and suppose that an event for a defective product was produced is denoted by by the machine . Then we have the following relations.


       , .

       


      


    이 정보를 가지고 전확률 공식을 사용하면, 불량품이 일어날 확률은

    ...

    입니다. 따라서 베이즈 정리에 의하여 불량품 중 기계 가 생산한 제품이 불량품일 확률은 다음과 같이 미리 구한 값들을 넣어 계산하면, 바로 4/33라는 답을 얻게 됩니다.


    Using the law of total probability with this information, the probability of defective products is obtained as follows.


              =  .  

      

    According to Bayes' theorem, the probability is 4/33.

        

    [열린문제 1] 베이즈 정리(Bayes’ theorem)가 적용되는 잘 알려진 문제가 조건부 확률에 관한 몬티홀(Monty Hall) 문제입니다.

    이 주소에서 https://ko.wikipedia.org/wiki/몬티_홀_문제 몬티홀 문제를 확인해보시고 몬티홀 문제에 대한 토론을 해주시면 충분히 조건부확률을 이해하실 수 있을 것입니다.

     

    그리고 이 강의를 녹화 촬영해 놓은 것이 이 주소 https://destrudo.tistory.com/5 에 있으니 자세하게 천천히 듣고 싶으면 이 주소를 활용하십시오.


    [Open Problem 1] Discuss the Monty Hall problem about conditional probability, which is a well-known problem that Bayes' theorem is applied.


     <https://en.wikipedia.org/wiki/몬티_홀_문제>

    Discuss the Monty Hall problem, and understand the conditional probability.

     

    Please refer to this video lecture https://destrudo.tistory.com/5.



    이어서 확률변수를 학습하겠습니다.

    동전 2개를 동시에 던져서 발생할 수 있는 사건은 4가지입니다.


    {앞면, 앞면},  {앞면, 뒷면},  {뒷면, 앞면},  {뒷면, 뒷면}


    이 각각의 사건이 일어나는 확률은 입니다. 이때 뒷면이 나오는 동전의 개수를 라 하면, 다음과 같이 각 사건은 숫자 0, 1, 2에 대응시킬 수 있습니다. 예를 들어, 은 (앞면, 앞면)에 대응시키고, (앞면, 뒷면)은 1로, (뒷면, 앞면)은 1로, (뒷면, 뒷면)은 2로 대응을 시킬 수가 있습니다. 우리가 이걸 표본공간으로 보고, 이걸 실수로 봐서, 함수 관계로 대응 시킨 겁니다.


       Next, we will cover the section 10.5. Random variables.


    In an experiment of flipping two coins, there are four possible outcomes:


    {Head, Head}, {Head, Tail}, {Tail, Head}, {Tail, Tail}


    The probability of each outcome is . If we set the number of tails in 2 tossess as , then each outcome corresponds to the numbers 0, 1, and 2 respectively. For example, we can match with (Head, Head) to 0, (Head, Tail) and (Tail, Head) to 1, and (Tail, Tail) to 2, as shown in the following figure. X is called a random variable.


    여기서 확률변수, random variable이란 컴퓨터 프로그래밍에서의 변수와 같은 것인데, 어떤 값을 취하느냐가 확률적으로 결정되는 변수입니다. 그래서 확률변수 또는 간단하게 '변수'라고 기술하기도 합니다.

    그 의미는, 표본 공간의 모든 표본에 대해 어떤 실수값을 대응시키는 것입니다. 따라서 확률변수를 사용하게 되면 구체적인 각 사건 대신에 이를 수치(數値)로 표현할 수 있어 여러 가지 계산과 분석이 가능해집니다. 확률변수는 영어 대문자로 쓰고 그 변수가 취할 수 있는 값 하나하나에 대해서는 소문자를 씁니다.


    A random variable works as a variable in computer programming. It is a variable whose value is determined probabilistically. So it is sometimes described as a random variable or simply 'a variable'.


    It means that some real value matches with all samples in sample space. Therefore, if random variable is used instead a probabilistic event, then it enables us to compute and analyze an random event. Random variables are written in uppercase letters and lowercase letters are used for each value that the variable can take.


    이번에는 확률분포에 대해 얘기합니다. 확률분포에는 이산확률분포와 연속확률분포가 있습니다. 다양한 예들이 있는데, 간단하게 소개하겠습니다.

    확률변수 가 연속적이지 않은 값 들을 취할 때, 를 이산확률변수라 하고, 각각의 에 대하여 일 확률 라고 할당한 것을 이산확률분포라고 합니다.


    Now, we will talk about <Section   10.6  probability distribution>. There are two types of probability distribution: <discrete probability distribution> and <continuous probability distribution.> There are various examples, and I will briefly introduce them.


    A discrete variable is a variable which can only take a countable number of values . And a discrete distribution describes the probability of occurrence of each value of a discrete random variable.


    다음과 같이 표로 나타낼 수 있는데, 확률이 , 가 일어날 확률이 , , 이 일어날 확률, 가 일어날 확률 그래서 전체가 일어날 합은 1이 될 것입니다.

    예를 들어서, 동전 2개를 동시에 던지는 시행에서, 뒷면이 나오는 동전의 개수의 확률 분포를 그림으로 나타내면 다음과 같습니다. (앞면, 앞면)이 나타날 확률변수를 0으로 대응시켰고, (앞면, 뒷면), (뒷면, 앞면)은 1로, (뒷면, 뒷면)은 2로 했으니까, 그러면 확률을 실제 생각해보면, (앞면, 앞면)이 일어날 확률은 1/4이 되고, (뒷면, 뒷면)이 일어날 확률은 1/4이 됩니다. 한 번은 앞면, 한 번은 뒷면이 일어날 확률은 1/2이 되는 것입니다. 두 가지 경우가 되어서 그렇습니다. 이때, 이산확률변수 가 값을 취할 때 확률 을 대응시키는 함수 를 확률변수 의 확률질량함수(probability mass function)이라 합니다.


    It can be expressed as a table as follows. and are the probabilities to the values and of the random variable , respectively. So the sum of the all probabilities is 1.


    For example, in an experiment of tossing two coins simultaneously, the probability distribution of the random variable defined by the number of tails is shown as the figure below. The probability of X = 0 is 1/4 since X = 0 corresponds (Head, Head). The probability of X = 1 is 1/2 since X = 1 corresponds (Head, Tail), (Tail, Head). The probability of X = 2 is 1/4 since X = 2 corresponds (Tail, Tail). 

    At this time, the function that matches the probability when the discrete random variable takes a value is called the probability mass function of the random variable.



    확률질량함수(probability mass function)은 가 주어졌을 때는 확률값이 존재하고, 그 외에 대해서는 0 값을 주는 함수입니다.

    이때, 다음의 3가지 성질을 만족합니다.

    The probability mass function is a function that gives a probability when is given, and a value of 0 otherwise.

    The probability mass function satisfies three properties as follows.


     ⓵   

     ⓶

     ⓷ when with  


    다음절이 연속확률분포인데, 연속확률분포에서는 적분의 지식이 필요하게 됩니다. 면적을 구해야하기 때문입니다.

    확률변수 가 어떤 범위에 속하는 모든 실수를 취할 때, 를 연속확률변수라고 합니다. 연속확률변수의 경우에는 특정한 값을 취할 확률 가 항상 0이므로, 확률질량함수를 이용하여 확률분포를 나타내는 것은 의미가 없습니다. 그래서 확률밀도함수(probability density function)를 새로 정의하여서 연속확률분포를 나타냅니다.

    연속확률변수 에 대하여 함수 가 다음 성질을 만족하면 를 의 확률밀도함수라 합니다. 확률질량함수의 성질에 대응되는 개념이라고 이해하시면 됩니다.


    The next section focuses on <the continuous probability distribution>. In the section on continuous probability distribution, we will study integration in order to determine the area.


    A continuous random variable is a random variable where the data can take infinitely many values. In the case of a continuous random variable, the probability of taking a specific value is always 0, so it makes no sense to use the probability mass function to represent the probability distribution. So, we should define a probability density function to represent the continuous probability distribution.


    For a continuous random variable , if the function satisfies the following properties, then is called the probability density function of . We can consider it as the concept that corresponds to the probability mass function of a discrete probability distribution.


     ⓵  for all real

     ⓶

     ⓷


    그림으로 보면, a에서 b 사이에 X가 있을 확률은 그림 안이 함수를 나타냅니다. 이 확률은 이 면적이 되고, 전체 면적을 구하면 total 확률이 1이 된다고 보시면 됩니다. a에서 b까지 일어날 확률은 여기를 구하면 되는 것입니다.


    여기서 ‘밀도(density)’라는 단어가 어떻게 쓰이게 되었을까? 생각해보면, 확률을 일종의 양(질량)으로 보고, 구간의 길이를 일종의 부피로 본다면, [확률/구간의 길이][구간의 길이] [확률]이므로 [확률/구간의 길이]는 [질량/부피]와 같은 의미가 되니까 여기서는 부피, 면적이라는 의미죠. 그러니까 '밀도'를 의미하게 되고, 이런 이유로 '확률밀도함수'라는 용어하게 된 것입니다.


    If we look at the graph of a function below, we can see that the probability of having an X between a and b represents this area. So the total probability, that is, the total area, must be 1.


    How did the word “density” come to be used here? If we think the probability as a kind of quantity (mass) and the length of the interval as a kind of volume, then we know [probability/length of the interval] is [mass/volume], from the relationship of [probability/length of the interval] * [length of the interval] = [probability]. So, it really means 'density', because of this, the term 'probability density function' is used,


    이상 전체 내용을 마쳤으니, 이제부터 복습을 해보겠습니다.

    이번 10주차에는 순열, 조합, 확률, 조건부 확률, 확률변수, 확률분포, 베이즈 정리, 이산확률분포와 연속확률분포에 대해서 학습을 했습니다.

    여기서 배운 기초통계지식은 인공지능 학습에 반드시 필요합니다.

    W10 주소에 실습실이 있으니까 여러분들이 직접 실습해보시면 됩니다.


    다음 시간에는 기댓값, 분산 variance, 공분산 covariance와 공분산 행렬에 대해서 학습하도록 하겠습니다.

    10주차 수고 많이 하셨습니다.


    We have finished the lecture of Week 10.


    Let’s review it.


    In this week, we learned about permutations, combinations, probabilities, conditional probabilities, random variables, probability distributions, Bayes’ theorem, discrete probability distributions, and continuous probability distributions.


    These basic statistical knowledges are essential for learning artificial intelligence.


    Please visit the link http://matrix.skku.ac.kr/math4ai-intro/W10/. In the next lecture, we will learn about expected value, variance, covariance, and covariance matrix of random variables.


    Thank you.


               그림입니다.
원본 그림의 이름: CLP0000422c0347.bmp
원본 그림의 크기: 가로 697pixel, 세로 522pixel







    Week 11. Introductory Mathematics for Artificial Intelligence

     

                인공지능을 위한 기초수학 입문


    11. 기댓값, 분산, 공분산, 상관계수, 공분산 행렬 122

    11.1 기댓값, 분산, 표준편차

    11.2 결합 확률분포

    *11.3 공분산, 상관계수

    11.4 공분산 행렬

      - 과제 (열린문제)-  133


    [11-1 pre]


    반갑습니다. K-MOOC Math4AI 11주차 강의입니다.


    Welcome!  We are now in Week 11.


    이번 주는 ‘기댓값, 분산, 공분산, 상관계수, 공분산 행렬’에 대해서 학습하도록 하겠습니다.


    In this week, we will study ‘expected value, variance, covariance, correlation coefficient, covariance matrix.’


    먼저 확률변수의 기댓값과 분산, 표준편차의 개념에 대하여 학습하고, 특히 2개 이상의 확률분포가 있을 때 확률분포간의 상관관계를 알 수 있는 공분산(covariance)에 대해서 학습합니다.


    First, we learn about the concepts of expected value, variance, and standard deviation of random variables, and especially covariance, which can know the correlation between random variables when there are two or more probability distributions.


    이 공분산의 개념을 이용하여 공분산 행렬(covariance matrix)를 만들고, 이 공분산 행렬을 활용하여 인공지능의 주요 개념인 주성분 분석(principal component analysis)에 바로 적용합니다.


    With the concept of covariance, a covariance matrix can be defined. This covariance matrix can be applied in the principal component analysis, one of the main concept in this course.


    순서는 1절에서는 기댓값, 분산, 표준편차에 대해서 학습하고, 2절에서는 결합 확률분포, 3절에서는 공분산과 상관계수, 그리고 4절에서는 공분산 행렬에 대해서 학습합니다.


    We learn about the expected value, variance, and standard deviation in Section 1, and the joint probability distribution in Section 2, the covariance and correlation coefficient in Section 3, and the covariance matrix in Section 4.


    먼저 기댓값 (뮤, ), 분산 (시그마 스퀘어, ), 표준편차 (시그마, )를 계산하는 과정을 배우고 또 코드를 학습하여, 데이터가 주어지면 언제나 뮤, 시그마를 쉽게 구할 수 있도록 실습 하고, 공분산 개념을 활용하여 공분산 행렬의 표준편차(시그마, )를 구할 수 있도록 학습할 것입니다.


    In the first session, we learn the notion of expected value (μ), variance (σ^2), standard deviation (σ), and covariance, and then we will use codes, so we can easily compute μ and σ from the data.


    [11주차 1강] http://matrix.skku.ac.kr/math4ai-intro/W11/


    반갑습니다. 

    고교생과 일반인을 위한 K-MOOC [Introductory Mathematics for Artificial Intelligence]. 11주차 강의를 시작하겠습니다.


    [Lec 1 Week 11] http://matrix.skku.ac.kr/math4ai-intro/W11/


    Hi everyone.  Let’s start our Week 11 of K-MOOC  “Introductory Mathematics for Artificial Intelligence”.


    11주차에는 기본적으로 분산, 공분산, 공분산 행렬을 중심으로 배우게 됩니다.


    In this week, you will basically learn <variance, covariance, and covariance matrices>.


    먼저 1절은 기댓값과 분산, 표준편차에 대해서 학습을 하겠습니다.


    First, in Section 1, we will study <an expected value, variance, and standard deviation>.


    그 실습실은 다음과 같이 (http://matrix.skku.ac.kr/math4ai-intro/W11/)에 준비되어 있으니, 주저 없이 실습실을 통해서 기댓값을 구하고, 분산을 구하고, 표준편차를 구하는 실습을 쉽게 하실 수 있습니다.


    The lab link http://matrix.skku.ac.kr/math4ai-intro/W11/  was made for you. You can easily find the expected value, variance, and standard deviation in R or Sage (programming language).


    먼저 기댓값으로 시작하겠습니다.

    확률변수의 기댓값(expectation)은 확률적 사건에 대한 평균값으로, 사건이 일어나서 얻는 값과 그 사건이 일어날 확률을 곱한 것을 모든 사건에 대해 합한 값입니다. 이것은 어떤 확률적 사건에 대한 평균을 의미합니다. 확률변수의 분산, variance은 그 확률변수가 기댓값으로부터 얼마나 떨어진 곳에 분포하는지를 가늠하는 수이고, 표준편차는 standard deviation이라고 부르죠, 분산의 양의 제곱근으로 정의됩니다.

     

    Let’s begin with the expected value. The expected value, or expectation, of a random variable is an average value for probabilistic events, which is the sum of products of a value obtained by each event and the probability of each event. The variance of a random variable is a measure of dispersion of numbers (data), which indicates how far a set of numbers (data) are spread out from their average value. The standard deviation is defined as a square root of variance.


      이산확률변수 의 기댓값과 분산, 표준편차는 다음과 같이 구할 수 있습니다.

    (화면의 수식을 보세요)


    You can find the expected value, variance, and standard deviation of a discrete random variable X with these formulas.

     (See the formulas on the screen)


      (1) 기댓값 : 

      (2) 분산 : 

      (3) 표준편차 : 

                                                             <---화면에 있으니 자막에는 없어도 됨)


    분산만 구하면 표준편차는 바로 구해지는 것이지요.

    You only need to find variance when you want to find the standard deviation.


    [예제 1] 확률변수 의 확률분포가 다음과 같을 때, 기댓값과 분산, 표준편차를 구하시오 하는 문제입니다. 0일 확률은 0.01이고 1일 확률은 0.84이고 2일 확률은 0.145고 3일 확률이 0.005, 그래서 전체 확률이 1라고 할 때, 다음과 같이 쉽게 Sage 코드를 활용해서 구할 수 있습니다. (화면의 코드를 보세요)


    [Example 1] Find the expected value, variance, and standard deviation of a random variable X. The probability of X being 0 is 0.01, 1 is 0.84, 2 is 0.145, and 3 is 0.005. Given that the total probability of an event is 1, this problem can be solved easily using the Sage code. (See the codes on the screen).


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    var_x = [0, 1, 2, 3]  # random variable

    prob_x = [0.010, 0.840, 0.145, 0.005]

    E_x = sum(var_x[i]*prob_x[i] for i in range(4)) # expected value

    V_x = sum((var_x[i] - E_x)^2*prob_x[i] for i in range(4))  # variance

    S_x = sqrt(V_x)  # standard deviation

    print("E(X) =", E_x.n(digits = 7))

    print("V(X) =", V_x.n(digits = 7))

    print("S(X) =", S_x.n(digits = 7))

    ---------------------------------------------------------------------

                                                             <---화면에 있으니 자막에는 없어도 됨)

    E(X) = 1.145000      # expected value  

    V(X) = 0.1539750      # variance     

    S(X) = 0.3923965     # standard deviation                        ■


    이것을 R 명령어, 통계학에서 많이 쓰이는 R 명령어로 구한다면 더 쉽습니다.

    (화면의 코드를 보세요)

       It becomes much easier to solve if you use R language, which is widely used for statistics.   (See the codes on the screen).

      

    ---------- http://matrix.skku.ac.kr/KOFAC/  ---  < R  code> ---------

    x <- c(0, 1, 2, 3)                      # 확률변수      < in R>

    pr.x <- c (0.010, 0.840, 0.145, 0.005)  # 확률분포

    e.x <- sum(x*pr.x)                      # 기대값

    var.x <- sum((x^2)*pr.x) - e.x^2       # 분산

    sd.x <- sqrt(var.x)                     # 표준편차

    cat(e.x, var.x, sd.x)                   # 한 번에 여러 개 출력

    ---------------------------------------------------------------------

                                                             <---화면에 있으니 자막에는 없어도 됨)

    1.145          # 기댓값   

    0.153975       # 분산      

    0.3923965      # 표준편차                                 ■


    필요에 따라서 원하는 명령어를 사용하시면 됩니다.

    We can decide which code to be used for our needs.


    연속확률변수 의 기댓값과 분산, 표준편차는 다음과 같이 계산합니다.

    (화면의 수식을 보세요)

       The expected value, variance, and standard deviation of a continuous random variable X can be found as follows:

       (See the formulas on the screen).

      (1) expected value : 

      (2) variance : 

      (3) standard deviation : 

                                                             <---화면에 있으니 자막에는 없어도 됨)


    이산확률변수에서는 이산값을 sum합니다. 연속확률변수에서는 함수가 연속함수이니까, 적분을 하게 됩니다. 이산확률변수와 연속확률변수의 기댓값과 분산, 표준편차는 똑같은 방식으로 구하는 거죠.


    When a random variable is discrete, we add up all discrete values. When a random variable is continuous, as values are not discrete, we integrate the function. So, the way to get the expected value, variance, and standard deviation is actually the same. 


    기댓값과 분산에 대해서 다음 성질이 만족합니다.

    (화면의 수식을 보세요)


       Here are the properties of the expected value and variance.

       (See the formulas on the screen) .

     

      ⓵ ,       

      ⓶ ,        

      ⓷ 확률변수 에 대하여 새로운 확률변수 를 로 정의하면, 위의 성질 ⓵, ⓶에 의해 평균과 분산이 항상 과 이다. 따라서 이 확률변수 를 확률변수 의 표준화 확률변수(standardized random variable)라 한다.

                                                             <---화면에 있으니 자막에는 없어도 됨)


    ⓵은 선형성을 보존을 뜻합니다. 그리고 이 성질들을 이용하면 여러 가지 형태의 기댓값과 분산을 쉽게 구할 수 있습니다. 이때, 중요하게 사용되는 개념이 표준화 확률변수인데, 확률변수 를 약간 변형시켜서 를 라고 놓으면, 위의 성질 ⓵, ⓶에 의해 평균은 0, 분산은 항상 1이 됩니다. 따라서 이 확률변수 를 확률변수 의 표준화 확률변수standardized random variable라고 합니다. 그러면 같은 확률변수를 분석할 때, 모든 계산이 아주 간단해 집니다. 그리고 원래 문제에 답을 주는 것도 어렵지 않습니다. 


    ⓵ is a linearity property. These properties make it easier to find the expected value and variance. One important concept is the standardized random variable, which is Z=. According to the properties ⓵ and ⓶ above, the expected value of Z is always 0, and the variance of Z is always 1. Thus, we call this variable Z a standardized random variable of X. Z makes the calculation simple. Also, it is not difficult to find answer for the original problem.


    [예제 2] 확률변수 의 확률밀도함수가 일 때, 와 의 분산을 구해보십시다. (화면의 수식을 보세요)


    [Example 2] Let the probability density function of variable is . Find the variance of and . (See the formulas on the screen).


    Solution.    

          

         

          

           ■

                                                             <---화면에 있으니 자막에는 없어도 됨)


    이 문제를 풀려면 적분지식이 필요합니다. 그래서 적분을 배웠던 겁니다.

    You need to know how to integrate a function if you want to solve this problem. That’s why we learned <Integral Calculus> in high school.


    확률밀도함수들이 변할 때마다 이 과정을 모두 반복해서 해야 되니까 코드로 짜주면, 파이썬 기반의 Sage 코드로 짜주면, 다음과 같이 만들 수 있습니다. (화면의 코드를 보세요)


    Whenever the function is changed, we may have to do this over again. So we may write a code for it in Sage. It is a computer language based on Python. (See the codes on the screen).


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    var('t')

    f(t) = 3*t^2

    Ex = integral(t*f(t), t, 0, 1)

    Ex2 = integral(t^2*f(t), t, 0, 1)

    print(Ex)                  # mean of X

    print(Ex2 - Ex^2)          #  variance of X

    Ey = integral((4*t + 2)*f(t), t, 0, 1)

    Ey2 = integral((4*t + 2)^2*f(t), t, 0, 1)

    print(Ey)                 # mean of Y = 4X + 2

    print(Ey2 - Ey^2)        # variance of Y

    ---------------------------------------------------------------------

                                                             <---화면에 있으니 자막에는 없어도 됨)


    3/4

    3/80    # variance of X,

    5

    3/5 # variance of ,    ■


    이 값이 앞에서 우리가 손으로 구한 값과 정확히 일치하는 것을 확인할 수 있죠. 그러니까 앞으로는 함수만 정의해주면 이 코드를 그대로 사용하셔서 기댓값, 분산, 표준편차를 구하실 수가 있을 겁니다.


    Now we can see that exactly the same answers can be found easily. So, we now only have to define a function differently (when it is needed) and use the same 'Code' from now on. We will get the expected value, variance, and standard deviation whenever we need without any difficulty.


    [열린문제 2] 연습문제로, 여러분이 다른 교재에서 찾은 연속확률변수의 기댓값, 분산, 표준편차를 직접 구해보시면 되겠습니다.


    [Open Problem 2] Try to find a continuous random variable X from other textbooks for your exercise. Find the expected value, variance, and standard deviation of X.


    이어서 결합확률분포. 확률변수가 두 개 이상 있는 경우에 각각의 확률변수에 대한 확률분포 이외에도 확률분포 쌍이 가지는 복합적인 확률분포를 분석할 필요가 있습니다.


    Now, let’s move on to <Joint Probability Distribution (Section 2)>. When there are two or more random variables, we should also look at the probability distribution for those random variables as well as the probability distributions for each variable.


    두 확률변수 값의 쌍이 어떤 확률분포를 가지는지 안다면 둘 중 하나의 확률분포의 값을 알고 있을 때 다른 확률분포가 어떻게 변하는지 알 수 있습니다. 이를 위하여 결합확률분포(joint probability distribution)에 대한 개념이 필요하게 된 겁니다.

     

    When we know the probability distribution for two variables, we can see how the probability distribution for one random variable changes depending on that of the other variable, only with the knowledge of one variable’s probability distribution. That’s the reason we need the concept of joint probability distribution.


    먼저, 이산확률변수인 경우를 살펴보겠습니다.

    와 가 이산확률변수일 때, 와 의 결합 확률함수(joint probability function)는 다음과 같이 정의합니다.

     

    Let’s look at the case of two being discrete random variables.

    If and are discrete random variables, the joint probability function for and is defined as below:


    이 와 의 가능한 모든 값에 대하여 의 값을 나타낸 것이 결합확률분포, joint probability distribution입니다. 이걸 표로 나타내면 다음과 같습니다. (화면의 표를 보세요)


    This is the joint probability distribution which shows for every possible value of  and . You can also show this in a chart. (See the chart on the screen).


     

         

     

     

    Sum

     

    marginal probability distribution for

     

     

     

    Sum

     

     

    marginal probability distribution for

     

     

     

     

                                                             <---화면에 있으니 자막에는 없어도 됨)


    이 표가 의미하는 바를 분석해보겠습니다.

    먼저 , , 들을 쭉 합친 것이 입니다. 일 확률 전체가 되는 것이지요. 그 다음 는 , , 를 다 더한 것입니다. 이것은 , , 을 다 더한 것입니다. 이 각각을 에 관한 주변 확률 분포(marginal probability distribution)이라고 부릅니다. 같은 방식으로 에 대해서 , , 을 더한 것을 이것이 에 관한 주변확률분포이고, 에 대해서도 주변확률분포를 차근차근히 구해서 에 대한 주변확률분포를 같은 방식으로 열(column)을 더해서 얻어질 수 있습니다.

    이것들을 다 더하면 total probability는 1이 됩니다. 행(row)을 다 더했을 때, 또 열(column)을 다 더했을 때도 확률은 1이 됩니다.


    Let me tell you what this chart indicates.


    First, is the sum of , , …, and , which is the total probability of . We call each one of the marginal probability distribution of . In the same way, is the sum of , , …, and and they are the marginal probability distributions of .

    When we add up all, whether we add all rows or columns, the result (total probability) is 1.


    , 의 결합분포가 주어져 있을 때 주변확률분포는 다음과 같이 정의됩니다.


    When the joint distribution for and was given, the marginal probability distribution is defined as below.


      

      


    이때, 는 행의 성분들의 합(row sum)이 되고, 는 열의 성분들의 합(column sum)이 되는 겁니다.

    즉, 주변확률분포란 결합확률분포에서 하나의 확률변수만 고려한 확률분포를 의미합니다.


    Here, is the row sum, and is the column sum.


    That is, the marginal probability distribution is the probability distribution for one variable from the joint probability distribution.


    결합확률분포에 관하여 아래와 같은 성질이 성립합니다.


    Here are the properties of the joint probability distribution. See the condition below:


     ⓵    , for all  with

     ⓶  , for all 

     ⓷  . for all 


    [예제 3]에서 크기가 같은 파란 색 공 개와 붉은 색 공 개, 녹색 공 개가 한 주머니에 들어 있을 때, 이 주머니에서 임의로 개의 공을 꺼낼 경우를 생각해보겠습니다. 꺼낸 공 중에서 파란색 공의 수를 , 붉은 색 공의 개수를 라 할 때, 와 의 결합 확률함수를 구하여라. 두 번째, 결합확률분포를 작성하여라. 그 다음 을 구하라. 의 주변분포를 구하라. 의 주변분포를 구하라. 이렇게 다섯 가지 문제가 주어졌습니다. 이걸 어떻게 계산하면 될지 생각해보죠.


    In [Example 3], it says there are three blue balls, two red balls, and three green balls of the same size in a pocket and we are going to take out two balls randomly. Let’s say that among the balls we took out, the number of blue ball is , and the number of red ball is . The problem is (1) Find the joint probability function of and , (2) Write down the joint probability distribution, (3) Find , (4) Find the marginal distribution of , and (5) Find the marginal distribution of .


     Now we are given these five problems. Let’s think about how we can solve them.


    풀이. 결합분포, 주변분포의 정의에 의해 각각을 계산하면 다음과 같습니다.

     

    Solution. We can get the chart below when we calculate each value, using the definition of joint distribution and marginal distribution.


     

     

         

    Sum

    합

     


     

    합

     


     

    Sum

     



    이 문제를 푸는데 표를 만드는 게 얼마나 중요한지 알 수 있었겠죠.


    Now we can see how important it is to make a chart to solve this problem.


    연속확률변수에 대하여 다음과 같은 결합밀도함수를 이용하여 결합확률분포를 나타내는데, 연속확률변수 와 의 결합밀도함수(joint density function) 는 다음과 같이 정의됩니다.


    As for a continuous random variable, we use the joint density function. The joint density function for continuous random variable and is defined as below.


     ⓵  , for all real

     ⓶  , for all real

     ⓷  , for all real

     ⓸ 가 평면상의 임의 영역 에 들어갈 확률은 다음과 같이 주어진다.


                       

     ⓹ 와 의 Marginal probability density function는 각각 다음과 같이 정의된다.


                  , 


    여기서는 변수가 2개가 되니까 다중적분, 이중적분, 삼중적분 등등이 필요하게 되는데, 다중적분에 대해서 필요한 경우에는 https://youtu.be/T1z_GYt85rI 에서 중적분을 제가 강의한 걸 참고하시면 되겠습니다.


    Since there are two variables, we need a multiple integral. So just like this problem, when you have to use the multiple integral, but you don’t know how to integrate, the following video lecture of mine will help you, https://youtu.be/T1z_GYt85rI.


    [예제 4] 두 확률변수 , 의 결합밀도함수가 다음과 같이 주어져 있을 때, 이 영역에서 와 의 주변밀도함수 이걸 구해보라는 문제입니다. 정의에 의해 구하면 다음과 같습니다.


    [Example 4] The joint density function of the two random variable and is given as below. Find the marginal density function of , in the given area. Using the definition, we can solve this problem as follows.


       ,  .

          , 


    확률밀도함수가 주어질 때마다 주변밀도함수를 수시로 구해야 하기 때문에 코드를 짜놓았습니다. (화면의 코드를 보세요)


    Whenever the probability density function is changed, we have to find the marginal density function over again, so I wrote a <Code> for it. (See the codes on the screen).


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    var( 'x, y' )

    f( x, y ) = x + y

    print( "f_X(x) =" , integral( f(x, y), y, 0, 1))  # X의 주변밀도함수

    print( "f_Y(y) =" , integral( f(x, y), x, 0, 1))  # Y의 주변밀도함수

    ---------------------------------------------------------------------

    f_X(x) = x + 1/2     # X의 주변밀도함수

    f_Y(y) = y + 1/2     # Y의 주변밀도함수                        ■


                                                             <---화면에 있으니 자막에는 없어도 됨)


    이상에서, 확률의 중요한 부분을 일단 마치고 이어서 다음 시간에는 공분산, 공분산 행렬을 학습하도록 하겠습니다.  수고 많이 하셨습니다.



    That’s all for this video. In the next video, you will learn covariance and covariance matrix.  Thank you.


    [11-2 pre]


    이제 11주 2차시는 ‘공분산 행렬’을 배우는 시간입니다.


    Hi!!  Today, we learn the 'covariance matrix'.


    앞에서 ‘기댓값, 분산, 공분산’에 대해서 학습을 했습니다.


    In the previous session, we learn about 'expected value, variance, and covariance'.


    이번 시간에는 ‘상관계수와 공분산’을 이용하여 ‘공분산 행렬’을 생성해 내는 과정을 이론뿐만 아니라 코드를 이용해서 실제 실습을 합니다.


    In this second session, we will practice the process of generating a ‘covariance matrix’ using ‘correlation coefficient and covariance’ with codes.


    여러분들은 주어진 데이터로부터 공분산 행렬을 생성하는 과정을 학습합니다.


    We learn the process of generating a covariance matrix from a given data.




    [11주차 2강] http://matrix.skku.ac.kr/math4ai-intro/W11/


    반갑습니다. 11주차 2차시입니다.

    오늘은 공분산과 공분산 행렬에 대해서 학습을 하겠습니다.

    그런데, 별표시(*)가 붙어 있는 내용이 있습니다. *표시는 고등학생에게는 어려운 내용일 수 있고 일반인들도 안 배웠거나 잊어버린 생소한 개념일 수 있음을 표시한 것입니다. 별표시 내용은 부담 없이 이해하시고 다음 절의 공분산 행렬에 대해서만 정확히 아시면 될 것 같습니다.


    [Lec 2 Week 11] http://matrix.skku.ac.kr/math4ai-intro/W11/


    Welcome. Hello. This is the second lecture of Week 11.

    Today, we are going to learn <Covariance and Covariance matrix>.


    Note: * Refers to optional.


    When we see this star symbol (*), it means that the part is a little bit more difficult and students are probably not ready yet. So, you don’t need to have any pressure at all. Just enjoy reading. You only need to know about <the Covariance matrix> to understand the Principal Component Analysis  in Chapter 5.


    3절, 공분산과 상관계수에 대해서 학습하겠습니다.

    하나의 확률변수 가 갖는 분포를 이해하기 위해서는 제일 먼저 사용하는 것이 평균입니다. 평균을 이용하면 분포에 관한 정보를 하나의 숫자, 분포의 중간 부분으로 나타낼 수 있습니다. 그리고 두 번째로 사용하는 개념이 바로 분산입니다. 분산을 이용하면 분포가 평균으로부터 얼마나 퍼져있는지를 알 수 있습니다.


    In Section 3, we will learn covariance and correlation coefficient.


    The mean is what we find first to understand the distribution of a random variable . It is a number that gives information about the distribution. And what we do next is to find the variance. The variance will help us to see how far a set of numbers are spread out from the mean.


    그렇다면 확률변수가 2개일 때, 이 확률분포들이 어떤 모양으로 되어있는지를 어떻게 알 수 있는가? 답은, 가장 먼저 의 평균과 의 평균을 생각할 수 있습니다. 그 다음으로 분산을 이용하여 각 확률변수가 얼마나 퍼져있는지를 알 수 있는데, 이때, 확률변수 간의 상관관계를 알기 위해서 공분산(covariance)의 개념을 사용하는 것입니다.


    Then how do we know what a probability distribution looks like when there are multiple random variables? The answer is,

    (1) First, we consider the mean of and .

    (2) Next, we will see the distribution of each variables using the variance. For this, we need the concept of <the Covariance>.


     확률변수 와 의 공분산은 다음과 같이 정의됩니다.


     The covariance of the random variable and is defined as follows:


       =


    이것은 기댓값의 성질에 따라서 로 쓸 수 있습니다.

    공분산 covariance는 의 편차와 의 편차를 곱한 것의 평균인데, 와 의 단위의 크기에 영향을 받는다는 문제점이 있습니다. 이것을 보완하기 위해 상관계수, correlation을 사용합니다.


    We can write this as because of the properties of the expected value. The covariance is the mean of the products of ’s deviation and ’s deviation. The problem is that the covariance is effected from the size (norm) of X and Y. That why we use the correlation.


    상관계수란 확률변수의 절대 크기에 영향을 받지 않도록 각 확률변수의 표준편차로 나누어서 표준화 시킨 것이라고 생각하면 됩니다.


    The correlation is a standardized value where we divide each deviation of and by their standard deviations before we calculate the products.


    확률변수 와 사이의 상관계수는 다음과 같이 정의됩니다. (화면의 수식을 보세요)


    The correlation of the random variable and is defined as follows. (See the formula on the screen).


                                                             <---화면에 있으니 자막에는 없어도 됨)



    이번에는 상관계수를 이용해서 공분산 행렬 covariance matrix을 정의하겠습니다.


    Now, we will define the covariance matrix using the correlation.


    행렬을 이용하면 여러 개의 확률변수가 서로 어떤 관계를 가지는지를 쉽게 표현할 수 있는데, 각 데이터의 분산과 공분산을 이용해 만드는 공분산 행렬이 바로 그것입니다.

    개의 확률변수 {, , }에 대한 공분산 행렬, covariance matrix는, 행렬의 성분이 일 때는 번째 확률변수 와 번째 확률변수 사이의 공분산 으로, 일 때는 번째 확률변수의 분산 으로 갖는 행렬로 정의하고 이 공분산을 로 표기합니다. 그러면 가 다음과 같이 주어집니다. (화면의 수식을 보세요)


    Using a matrix, we can easily see how multiple variables relate to one another. So, we are going to make a covariance matrix using the variance and covariance of each data set.


    The covariance matrix of 'p' random variables {, , } has the covariance between and () as its (i,j)-th entry when , while it has the variance of i-th random variable () as its (i,j)-th entry when . When , the matrix is defined as a matrix and written as . This is given as follows. (See the formula on the screen).


     

                                                             <---화면에 있으니 자막에는 없어도 됨)


    주대각선 성분은 , , 로 주어지고, 주대각선 성분 외에는 으로 다 구해질 수 있습니다. 그러면 더 간단하게 표시하면 주대각선 성분에는 이 있고, off-diagonal 성분에는 가 쭉 놓이게 됩니다. 이게 공분산 행렬입니다.


    The main diagonals are given as , , ... , and the other entries can be found using . That is, the main diagonals are   and the off-diagonals are . This is <the Covariance Matrix>.


    쉽게 말하면, 정사각행렬의 성분을 각 변수의 분산과 공분산으로 채운 것이 바로 공분산 행렬입니다. 각 변수의 분산을 주대각선 성분으로, off-diagonal 성분을 공분산으로 채운 행렬이 이 공분산 행렬입니다.


    Simply put, the covariance matrix is a square matrix whose entries are variance and covariance of each variables, where the variances are the main diagonals and the covariances are the off-diagonal entries.


    공분산 행렬은 아래 그림과 같이 데이터 분포를 나타나낸다고 볼 수 있습니다.

    이렇게 공분산 행렬을 보면, 데이터가 어떻게 변환되어 있는지를 이해할 수 있습니다.


    We can say that the covariance matrix shows data distribution as the figure below.

    Looking at these covariance matrices shows that how the distribution of data can be displayed.


    [예제 5] 표본 자료가 다음과 같이 주어져 있을 때 이 표본 자료로부터 (표본) 공분산 행렬을 구하라 즉, 이렇게 데이터가 있을 때, 공분산 행렬을 구하라는 문제입니다. 그러면 이제 파이썬 기반의 Sage 코드와 R코드를 이용하여서 공분산 행렬을 구해봅시다.  (화면의 코드를 보세요)


     [Example 5] Find the covariance matrix of the given sample data. This is the problem that asks us to <Find the Covariance Matrix from the given data>. Now, let’s use Sage (which is based on Python) and R language to find the covariance matrix. (See the codes on the screen).


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    X = matrix([[1, 2, 3, 4, 5, 6],

                [2, 3, 5, 6, 1, 9],

                [3, 5, 5, 5, 10, 8],

                [10, 20, 30, 40, 50, 55],

                [7, 8, 9, 4, 6, 10]])

    n, p = X.nrows(), X.ncols()  # n은 행의 개수, p는 열의 개수

    X_ctr = zero_matrix(RDF, n, p)  # 행렬 준비

    import numpy as np

    for row in range(n):

        v = np.array(X.row(row))

        m = mean(v)

        for col in range(p):

             X_ctr[row, col] = X[row, col] - m


    # 공분산 행렬

    print("Covariance_Matrix =")

    print(1/(p - 1)*X_ctr*X_ctr.transpose().n(digits = 7))

    ---------------------------------------------------------------------

    Covariance_Matrix =                                     # 공분산 행렬

    [ 3.500000  3.000000  4.000000  32.50000 0.4000000]

    [ 3.000000  8.666667 0.4000000  25.33333  2.466667]

    [ 4.000000 0.4000000  6.400000  38.00000 0.4000000]

    [ 32.50000  25.33333  38.00000  304.1667  1.333333]

    [0.4000000  2.466667 0.4000000  1.333333  4.666667]        ■


                                                             <---화면에 있으니 자막에는 없어도 됨)


    (화면의 코드를 보세요)

    공분산 행렬을 생성하는 R 코드 실습 - http://matrix.skku.ac.kr/KOFAC/ < R  명령어>

    참고 : https://stats.seandolinar.com/making-a-covariance-matrix-in-r/


    # 표본 자료 create vectors                 < R  명령어>

    a <- c(1,2,3,4,5,6)

    b <- c(2,3,5,6,1,9)

    c <- c(3,5,5,5,10,8)

    d <- c(10,20,30,40,50,55)

    e <- c(7,8,9,4,6,10)


    # create matrix from vectors

    M <- cbind(a,b,c,d,e)

    cov(M)                                        # 공분산 행렬

                                                             <---화면에 있으니 자막에는 없어도 됨)


    참고로, 공분산 행렬은 고차원 데이터의 분포를 최대한 유지하면서 차원을 효과적으로 줄이는 차원축소(dimension reduction)에서 굉장히 중요한 역할을 합니다. 대표적인 기법이 주성분 분석(principal component analysis, PCA)가 있으며, 이 주성분을 계산할 때 제일 앞에서 배운 특잇값 분해(singular value decomposition)을 사용합니다. 지금 covariance 행렬을 만들어 내는 것을 미적분, 다중적분을 이용해서 통계지식을 이용해서 공분산 행렬을 만들어냈습니다. 공분산 행렬을 인공지능에 활용하기 위해서 앞에 선형대수학 부분에서 배운 특잇값 분해와 앞으로 배울 주성분 분석을 이용해서 행렬의 차원을 축소해서 인공지능에 바로 활용하는 것입니다.


    For reference, the covariance matrix is very important in the dimension reduction. The most important technique used in dimension reduction is Principal Component Analysis (PCA). And we use Singular Value Decomposition (SVD) to compute <Principal Components>. We will use Calculus and Statistical knowledge to find out what covariance matrix is. We can use this covariance matrix in AI by reducing the dimension of the matrix. In Chapter 5, we will this matrix in PCA with SVD which we learned at the part of Linear Algebra.


    이상 수고하셨습니다.  오늘 배운, 11주차에 배운 내용을 간단히 요약하면, 다음과 같습니다.

    That’s all for today. Now let’s summarize what we have learned in this week.


    이번 11주차에는 통계의 기댓값, 분산, 공분산, 상관계수, 공분산 행렬(Covariance Matrix)에 대해서 학습했습니다. 내용은 분산, 공분산, 공분산 행렬을 중심에서 배웠고, 그것을 계산하기 위해서 기댓값, 분산, 표준편차 개념을 복습하였으며, 그것으로부터 결합확률분포를 배웠고, 그런 지식을 활용하여 공분산과 상관계수를 구하는 법을 배운 후에 공분산 행렬을 정의했습니다. 기댓값의 정의는 E(X), Expectation, 분산은 V(X), 우리가 로 쓰고, 그것에 를 씌워준 가 바로 표준편차 S(X)입니다. 그리고 Covariance 행렬은 다음과 근사한, 주대각선 성분은 variance로, off-diagonal 성분은 Covariance로 갖는 행렬입니다.


    In Week 11, we have learned the expected value, variance, covariance, correlation, and covariance matrix. We first reviewed the expected value, variance, and standard deviation. And then, we learned the joint probability distribution. After that, we learned covariance and correlation, and finally, we defined the covariance matrix.


    E(X) is the expected value, V(X), or , is the variance, and the square root of (), is the standard deviation, S(X). The covariance matrix is the matrix that has variances as its main diagonals and covariances as its off-diagonals.


    다음 시간 12주차에는 Principal component analysis를 시작하면서 바로 인공지능의 심장으로 들어가도록 하겠습니다.


    Next week, we will start learning the principal component analysis, which is the key concept of the Artificial Intelligence.


    수고 많으셨습니다.


    Hope you did enjoy this lecture. Thank you.




      그림입니다.
원본 그림의 이름: 12100469-1.jpg
원본 그림의 크기: 가로 5265pixel, 세로 973pixel
사진 찍은 날짜: 2019년 12월 10일, 오후 13:16

     

    Week 12. Introductory Mathematics for Artificial Intelligence

     

                인공지능을 위한 기초수학 입문

     


    V.  주성분 분석과 인공신경망 (12-14주차)   134

    12. 주성분 분석(Principal Component Analysis) 134

    12.1 차원 축소

    12.2 주성분 분석(PCA)

    12.3 주성분 분석의 계산

    12.4 주성분 분석 사례

    *12.5 주성분 분석과 공분산 행렬

    *12.6 주성분 분석과 선형회귀


    [12-1 pre]


    여러분 반갑습니다. K-MOOC Math4AI 12주차 에서는 주성분 분석(principal component analysis)을 학습합니다. 


    Welcome! In Week 12, we learn the principal component analysis.


    이번 주에 다룰 <주성분 분석 (PCA)>은 ‘인공지능 수학’에서 가장 중요한 부분이고, 2-3-4장에서 배운 모든 지식이 사용됩니다.


    <Principal Components Analysis (PCA)> is the most important part of this course.  And all the knowledge learned in Chapters 2-3-4 will be used here.


    대용량 또는 고차원 데이터는 계산과 시각화가 어려워 분석하기가 쉽지 않습니다. 따라서 원데이터의 분포를 가능한 유지하면서 데이터의 차원을 줄이는 과정이 필요합니다. 이를 차원 축소(dimension reduction)이라고 합니다.


    Large or high-dimensional data is difficult to analyze because it is difficult to compute and visualize. Therefore, it is necessary to reduce the dimension of the data while maintaining the distribution of the original data as much as possible. This is called a dimension reduction.


    주성분 분석은 가장 널리 사용되는 차원 축소 기법 중 하나로, 원 데이터의 분포를 최대한 보존하면서 고차원 공간의 데이터들을 저차원 공간 안에서, 우리가 다룰 수 있는 데이터로 변환시켜 줍니다.


    Principal component analysis is one of the most widely used a dimension reduction technique. It converts a data in high-dimensional space into a data that can be handled in a low-dimensional space while preserving the distribution of the original data as much as possible.


    주성분 분석을 학습하는 단계는 먼저 1절에서 차원축소를 배우고, PCA 개념을 이해한 후에, PCA 관련한 실제 계산을 해보고, 사례를 살펴본 후에, 주성분 분석이 공분산 행렬에 적용되는 과정을 학습할 것입니다.


    We first learn the concept of dimensionality reduction in Section 1, understand how to compute PCA through some examples, and learn relationship between principal components analysis and the covariance matrix.


    이렇게 배운 주성분 분석이 선형회귀에 어떻게 적용되는지 실습을 통해서 학습하도록 하겠습니다.


    Also, we compare the principal component analysis and the linear regression.


    데이터들이 흩어져있을 때 이로부터 차원축소를 위해 필요로 한 첫 번째 축,  두번째 축을 순서대로 구해가는 과정을 학습하고, 그리고 그를 위해서 앞에서 배운 SVD(Singular Value Decomposition)가 어떻게 활용되는지를 구체적으로 본 후, 우리가 새로 구한 주축(principal axis)이 어떤 영향을, 우리가 순서대로 찾은 주축(principal axis)과 PC(principal component)가 그리고 PC score가 앞에서 몇 개까지 있으면, 그 전체 behavior의 80%, 90% 또는 98% 내용을 설명할 수 있다는 것을 이해하고, SVD에서 얻은 특잇값(singular value)과 특이벡터(singular 벡터)들 몇 개만 골라서 작은 사이즈의 행렬을 만들어서 분석을 하고 거기서 나온 결과를 원래 문제로 돌아가서 적용 시키면서 우리가 원하는 의사결정을 할 수 있도록 해주는 기계학습(Machine Learning)의 코어(core) 부분이 여기서 시작되는 것입니다. PCA(principal component analysis)는 인공지능 학습에서 가장 흥미있는 그리고 중요한 이론입니다. http://matrix.skku.ac.kr/math4ai-intro/W12/


    From a given dataset, we learn the process of obtaining the first new axis and the second new axis required for dimension reduction. And we use the SVD (Singular Value Decomposition) learned earlier to compute principal component and principal axis.


    If the first few principal components explained 90% of variance, then we can only .\PICk these principal components to do data analytics techniques. This is why the PCA is the most fundamental and important part in AI.

    http://matrix.skku.ac.kr/math4ai-intro/W12/ 


    [12주차 1강]  반갑습니다. K-MOOC [인공지능을 위한 기초수학 입문].


    [Week 12, 1st lecture] http://matrix.skku.ac.kr/math4ai-intro/W12/

     

    Welcome to K-MOOC [Introduction to Basic Mathematics for Artificial Intelligence] class.


    그 내용의 핵심중의 핵심인 <주성분 분석과 인공신경망> Part5를 시작하도록 하겠습니다.


    Let's start with Part 5, <Principal Component Analysis and Artificial Neural Network>, which is one of the core to.\PICs of the course.


    이번 시간에는 <주성분 분석>에 대해서 학습합니다.


    In this lesson, we will study <Principal Components Analysis>.

     

    그 내용은 12주차 실습실 http://matrix.skku.ac.kr/math4ai-intro/W12/  에서 여러분들이 직접 실습해 보시게 될 것이고, 우리도 다음 차시에 실습을 하면서 그 내용을 실제 구현해 볼 것입니다.


    We can practice the content of the week 12 practice room in the following lab link http://matrix.skku.ac.kr/math4ai-intro/W12/. We will practice it after we study the contents first.


    1절은 차원축소(dimension reduction)로 시작합니다.


    Lecture 1 of Week 12 begins with a “Dimension reduction”.


    다음은 1974년 Motor Trend magazine에 실린 1973년~74년도에 출시된 자동차의 특징을 기술하고 정량화하는 11가지 변수에 관한 데이터의 일부입니다.


    Here are some of the data on the 11 variables that characterize and quantify cars released in 1973-74, published in the <Motor Trend magazine> in 1974.


     자동차 Mazda RX4의 연비(Miles per gallon), 실린더 개수(몇 기통), 마력(horse power), 무게(weight) 등 11개의 변수들로 데이터를 수집하고, Mazda RX4 웨곤, Datsun, Hornet 4 Drive, Hornet Sportabout, Valiant 등의 차들을 쭉 나열한 것입니다. 각각의 연비, 마력, 실린더 개수 등의 이 데이터들은 이 주고 http://www.people.vcu.edu/~jpbrooks/pcatutorial/ 를 클릭하시면 바로 얻을 수 있습니다.


    From the fuel consumption of the Mazda RX4 car (miles per gallon), the number of cylinders, horsepower, weight and other similar data are collected, which will be represented by 11 variables, Then a set of data is collected in the same way for such cars as Mazda RX4 station wagon, Datsun, Hornet 4 Drive, Hornet, Sportabout and Valiant, etc. and are placed in order. All of these data, such as fuel consumption, horsepower and number of cylinders, can be obtained from the following link: http://www.people.vcu.edu/~jpbrooks/pcatutorial/.


    이 표에서 보듯이 자동차 한 대에 관한 데이터는 11차원의 벡터로 표현된 것입니다. 이와 같은 고차원의 데이터는 계산과 시각화가 어려워서 분석하기가 쉽지 않습니다. 따라서 원 데이터의 분포를 가능한 유지하면서 데이터의 차원을 줄이는 것이 필요한데 이를 차원 축소(dimensionality reduction)라 합니다.


    As we can see in this table, the data for each car is presented as an 11-dimensional vector. Such a multidimensional data is difficult to analyze because it is not easy to visualize and compute. Therefore, it is necessary to reduce the dimension of the data, while preserving the distribution of the original data as much as possible. This process is called a dimension reduction.


    물론 필요한 일부 변수의 데이터만 뽑아서 사용할 수도 있지만, 변수들 사이에 어떤 밀접한 관계가 있는지는 미리 알 수가 없어서 이 경우에 이런 식으로 임의로 차량의 무게와 배기량만 선택해서 분석을 한다고 해도, 원래 충분히 조사된 데이터의 그 분포를 제대로 반영한다고 볼 수는 없으므로 어떤 특색(feature)나 요소를 중심으로 분석을 해야 되는지 주성분(principal component)을 선택하는 분석이 필요합니다. 그것이 바로 주성분 분석(principal component analysis, PCA)입니다.

    Of course, it is possible to extract and only use the data of some essential variables, but it is impossible to know in advance what a close relationship exists between the variables, so even if we select and analyze only the weight and displacement (engine/piston/cylinder) in this way, we may not be sure that it properly reflects the distribution of the originally sufficiently investigated data, and therefore, an analysis of the choice of the principal component is necessary to determine which function or element to analyze. This process is called the Principal Component Analysis (PCA).


    주성분 분석(PCA)는 가장 널리 사용되는 차원 축소 기법 중 하나로, 원 데이터의 분포를 최대한 보존하면서 고차원 공간의 데이터들을 다루기 쉬운 적당한 차원의 저차원 공간으로 변환하는 것입니다.


    The Principal Component Analysis (PCA) is one of the most widely used a dimension reduction techniques, which converts a data set from high-dimensional space into a low-dimensional, easy-to-handle spaces while preserving the distribution of the original data as much as possible.


    PCA는 기존의 변수를 조합하여 서로 연관성이 없는 새로운 변수 즉, 주성분, principal component를 찾아냅니다. 첫 번째 주성분 PC1이 원래 데이터의 분포를 가장 많이 보존하고, 두 번째 주성분 PC2가 그 다음으로 원래 데이터의 분포를 많이 보존하는 식입니다.


    The PCA combines existing variables to find new variables that are not related to each other, namely, principal components. The first principal component PC1 preserves the distribution of the original data as much as possible, the second principal component PC2 preserves the distribution of the next original data as much as possible too, and so on.


    앞서 언급한 11차원의 데이터의 경우 기존의 변수들을 조합하여 같은 개수(11개)의 주성분을 만들 수 있는데, 만일 PC1, PC2, PC3가 원 데이터의 분포(성질)의 약 90%를 보존한다면, 10% 정도의 정보는 잃어버리더라도, 합리적인 분석에 큰 무리가 없으므로, PC1, PC2, PC3만 택하여 3차원 데이터로 차원을 줄일 수 있습니다. 이 경우 계산과 시각화가 용이하여 데이터를 쉽게 분석할 수 있습니다.


    In the case of 11-dimensional data mentioned above, the same number (11) of principal components can be created by combining existing variables. If PC1, PC2, and PC3 preserve about 90% of the distribution (property) of the original data, even if about 10% of the information is lost, there is no big problem in a rational analysis, so we can simply reduce the dimension to 3D data by selecting only PC1, PC2, and PC3 for the next level of analysis. In this case, it is much easier to compute and visualize, so it make us to analyze the data without any difficulty.


    주성분을 구하기 위해서 새로운 축을 찾아야 되는데, 이를 주축(principal axes) 또는 주방향(principal direction)이라고 합니다. 첫 번째 축은 원래 데이터를 이 축 상에 정사영(projection) 하여 얻어진 데이터들의 분포가 가장 최대가 되도록 정합니다. 아래는 주어진 2차원 데이터(왼쪽 그림)를 첫 번째 축을 찾아 정사영한 그림(오른쪽 그림)입니다. 이때 정사영 된 데이터들이 PC1을 이루게 됩니다.


    To find the principal component, we need to find a new axis, which is called <the principal axes or principal direction>. The first axis sets the largest distribution of data obtained by projecting the original data onto this axis. Below is a .\PICture of the given 2D data (left .\PICture) orthogonal to the first axis (right .\PICture). At this time, the projected data constitutes PC1.


    데이터가 다음과 같이 주어져 있을 때는 분석을 어떻게 해야 될까요? 정사영을 이용하여 오차(error)가 상대적으로 적어지는 새로운 축을 찾는 것입니다. 그래서 첫 번째 축을 가장 먼저 잡고, 그 다음 직교인 두 번째 축을 잡고, 같은 방법으로 PC1, PC2, PC3를 잡는 것입니다.


    When the data are given like this, how do we analyze them?


    Orthogonal projection should be used to find the new axis where the error is relatively small. So, we first take the first axis, then the second orthogonal axis, and we continue to do this for  PC1, PC2 and PC3, etc in a same way.


    (센터링된) 컴퓨터과학 전공생 105명의 고등학교 성적(GPA)과 대학교 성적(GPA) 데이터

    는 [출처] http://onlinestatbook.com/2/regression/intro.html 와

    http://onlinestatbook.com/2/case_studies/sat.html 에 있습니다.


    [Source of the data]  High school grades (GPA) and university grades (GPA) data of 105 computer science students (centered) are available at the following links. Source:

    http://onlinestatbook.com/2/regression/intro.html

    http://onlinestatbook.com/2/case_studies/sat.html.


    두 번째 축은 어떻게 찾느냐? 두 번째 축도 마찬가지로 원래 데이터를 이 축 상에 정사영할 때, 얻어진 데이터들의 분포가 PC1 다음으로 가장 최대가 되게 합니다. 역시 정사영(projection)의 개념을 이용하여, 오차들의 합의 크기를 고려하는 것입니다. 그리고 PC1과 PC2가 서로 관계가 없도록, 즉 두 번째 축은 첫 번째 축과 관계가 없도록 수직이 되도록 하는 것입니다. 우리가 정사영을 배우고 Gram-Schmidt 정규직교화법을 배우는 게, 바로 이 개념을 이해하기 위해서입니다.


    How to find the second axis? Similarly for the second axis, when the original data is projected onto this axis, the distribution of the obtained data is the largest after PC1. Again, using the concept of projection, we consider the size of the sum of the errors. And so on, then PC1 and PC2 are not related to each other, that means the second axis is not related to the first axis. The reason that we did learn the orthogonal projection and Gram-Schmidt orthonormalization process was for this. Now you see the reason.


    그 이유는 다음과 같습니다. 앞서 차원 축소를 소개할 때 필요한 일부 변수의 데이터만 뽑을 수도 있다고 하였으나 이런 경우는 변수들 사이에 어떤 밀접한 관계가 있는지는 미리 알 수 없어서 변수 간에 어떤 영향을 미치는 지 파악할 수가 없습니다.


    The reason is as follows. Previously, when introducing a dimension reduction, it was said that it is also possible to extract only some variables that are needed from the origianl big data, but in this case, it is not possible to know in advance what kind of close relationship does exist between these variables, so it is not possible to understand the effect of the variables at the first stage.


    주성분 분석은 이를 방지하기 위해 관계가 있는 변수끼리는 각 주성분에 모이도록 하고, 주성분들 사이에는 서로 관계가 없도록, (가능하면) 일차독립이 되도록 축을 잡아 주는 것입니다. 이는 행렬의 대각화에 관계되는 개념입니다. 그래서 서로 다른 주축이 직교가 되도록 하는 것입니다. 정규직교화법에서 서로 일차독립인 축들을 찾아가는 과정이 바로 여기, 주성분분석에 활용되는 겁니다. 이를 위해서 선형대수학, 특히 행렬의 대각화를 배웠다고 이해하셔도 됩니다.


    To prevent this, a principal component analysis is needed to set the axis so that essential variables are collected as each principal components, and that each principal components are not related to each other and are (if possible) linearly independent. This is a concept related to the diagonalization of matrices. This is why it is necessary to ensure the orthogonality of the different principal axes. The process of finding axes that are linearly independent of each other in the orthogonalization method is used for principal component analysis. We could say that it was in order to be able to do this that you studied Linear Algebra, and in particular about the matrix diagonalization (or SVD).


    아래 왼쪽 그림은 주어진 2차원 데이터를 두 번째 축을 찾아 projection을 한 그림입니다. 이때 정사영 된 데이터들이 PC2을 이루게 됩니다. 그리고 첫 번째 주축을 축, 두 번째 주축을 축으로 놓고 데이터를 정사영 시키면 오른쪽 그림과 같이 되겠죠. x축 잡아주고, y축 잡아주고. 아까 x축을 이렇게 잡아줬습니다. 그리고 y축을 이렇게 잡아주고. 이로부터 첫 번째 주축을 따라 정사영 된 데이터의 분포가 훨씬 큼을 직관적으로 이해할 수 있습니다. x축 분포가 더 길고, y축 분포가 더 작습니다.


    The .\PICture on the left below is a projection of the given 2-dim data by finding the second axis. Here, the projected data forms PC2. And if you put the first main axis as the -axis, the second main axis as -axis and do project the data orthogonally, it will be looked like a .\PICture on the right. Hold the - and -axes like this. I took the -axis like this. And take the -axis like this. From this, we can intuitively understand that the distribution of the orthogonal data along the first principal axis is much larger. The -axis distribution is longer and the -axis distribution is shorter.


    이어서, 주성분 분석을 실제 계산해 보도록 하겠습니다.


    Next, let's actually compute the principal component analysis.


      를 의 데이터 행렬, data matrix라 하지요. 여기서 은 표본(sample)의 개수이고, 는 데이터의 특성(feature)을 나타내는 확률변수 {, , }의 개수라고 합시다. 통계학에서는 일반적으로 데이터 행렬 와 구분하기 위하여 확률변수를 대문자 {, , } 대신해서 간단하게 소문자 {, , } 를 씁니다. 그래서 여기서는 확률변수를 소문자 {, , }로 사용하도록 하겠습니다.


     is called an data matrix. Suppose that is the number of samples, and is the number of random variables {, , } representing the features of the data. In statistics, in general, we simply use lowercase notation {, , } instead of uppercase notation {, , } for random variables to distinguish it from the data matrix . So, here we will use a random variable in the lower case {, , }.


    데이터 행렬 의 성분 는 번째 표본의 번째 확률변수 에 대한 하나의 데이터를 의미하고, 열 는 확률변수 가 갖는 모든 데이터를 의미한다고 하겠습니다.  (화면의 수식을 보세요)


    Let's say that the component of the data matrix means one data set for the th random variable in the -th sample, and column, means all the data set of the random variable . (Look at the formula on the screen).

               그림입니다.
원본 그림의 이름: CLP00003d0c0001.bmp
원본 그림의 크기: 가로 564pixel, 세로 310pixel  

     (Delete when putting in subtitles)

         <---화면에 있으니 자막에는 없어도 됨) (It's on the screen, so it doesn't have to be in the subtitles)


    위의 행렬 를 확률변수의 평균이 0되도록 센터링(centering, 확률변수의 평균을 0으로 조정)합니다. 여러 가지 이유가 있는데 일단 계산이 편리해지고 분석이 굉장히 용이해집니다. 센터링된 행렬을 (틸다)라고 정의합시다. X라는 행렬이 주어지면 센터링을 해서 확률변수의 평균이 0이 되도록 하는 행렬 를 정의합니다.


    Center ("centering, adjust the mean of the random variable to zero") the matrix above so, that the mean of the random variable is zero. There are a number of reasons for this. Firstly, the computations become convenient, and the analysis becomes very simple and easy. Let's define the centered matrix as (tilda). Given a matrix named X, we define an matrix, centered so, that the mean of the random variables is zero.


    이를 위해서 ⓵ 의 각 열(하나의 확률변수가 갖는 데이터)의 평균을 구합니다. 확률변수 의 평균은 로 표기합니다. (화면의 수식을 보세요)


       To do this, average each column of ⓵ (the data of one random variable). The mean of the random variable is denoted by . (Look at the formula on the screen)


      , ,  .

                                                             <---화면에 있으니 자막에는 없어도 됨) (It's on the screen, so it doesn't have to be in the subtitles)


     다음 ⓶ 각 열별로 데이터에서 (열의) 평균을 뺍니다. 이 행렬을 센터링된 행렬 라 합니다.    (화면의 수식을 보세요)


    Then ⓶, subtract the mean (of the column) from the data for each column. This matrix is called the centered matrix . (Look at the formula on the screen).


                                                             <---화면에 있으니 자막에는 없어도 됨) (It's on the screen, so it doesn't have to be in the subtitles)



    이렇게 센터링된 행렬은 평균을 구해서 그것을 빼서 만들어진 행렬이라 mean-centered 행렬이라 부릅니다. 앞으로 우리가 뒤의 분석에서 얘기할 때는 주어진 데이터 행렬은 간단하게 이 절차를 거쳐서 만들어진 센터링 된 행렬, 즉 mean-centered 행렬이라고 가정을 하고, 이론을 전개하겠습니다.


    This centered matrix is called a mean-centered matrix because it is created by obtaining an average and subtracting it. In the future, when we talk about subsequent analysis, we will develop the theory by assuming that a given data matrix is ​​simply a centered matrix created by this procedure, that is, a mean-centered matrix.


     ⓷  이제 mean-centered 행렬인 의 특잇값 분해(SVD)를 구합니다.  (화면의 수식을 보세요)


    ⓷ Now, find the Singular Value Decomposition (SVD) of a mean-centered matrix . (Look at the formula on the screen).


       

          

                                                          


    여기서 , 는 직교행렬이고, 는 크기 순서대로 배열된 특잇값(singular value) 을 주대각선 성분으로 하는 대각선행렬 입니다.


    Here, and are orthogonal matrices, and is the diagonal matrix with a single value as the main diagonal component arranged in order of magnitude.


    여기서 재미있는 것은 p개의 singular value들만 0보다 크고 나머지는 0입니다. 따라서 곱한 후 결과는 가 되고 이는 U 행렬에서 첫 번째 p개의 column만 사용되었다는 뜻이고, 직교행렬 V에서도 p개의 column만 사용되었다는 뜻입니다.


    The interesting thing here is that only p singular values are greater than 0 and the rest are 0. Hence, after the multiplication, the result is , which means that only the first p columns were used in the matrix U, and only p columns were used in the orthogonal matrix V.


    ⓸  의 특잇값 분해에서 의 열벡터는 principal axes, 주축이 되고,


    ⓸ In the Singular Value Decomposition of X, the column vector of V becomes the principal axes.


    ⓹ 와 의 곱을 로 표시하면, 의 열벡터들이 원 데이터를 주축에 정사영하여 얻어진 주성분 점수, principal component score, PC score가 되는 것입니다.


    ⓹ If the product of and is expressed as , the column vectors of obtained by orthogonal projection of the original data to the principal axis, become <principal component scores (PC scores)>.


     ⓺ 이 다음에 특잇값 를 이용하여 계산한 은 번째 principal component의 분산이 되고 은 번째 principal component가 원 데이터의 분포를 보존하는 비율이 됩니다. 각 PC별로 원 데이터의 분포를 보존하는 비율을 계산해서, 몇 차원으로 데이터의 차원을 축소할지를 결정할 수 있습니다. 마지막이 결정적인 내용입니다.

    ⓺ After that, computed with singular values becomes the variance of -th principal components, and becomes the ratio at which the -th principal component preserves the distribution of the original data. By calculating the proportion that preserves the distribution of the original data for each PC, you can decide in which dimension to reduce the dimension of the data. The last is the decisive stage.


    ⓻ 데이터를 차원에서 () 차원으로 줄이기 위하여, 의 처음 개의 열벡터()와 의 번째 선행 주 부분행렬(leading principal submatrix) ()을 택하면, 는 처음 개의 PC(주성분)를 포함하는 행렬이 됩니다.  (화면의 수식을 보세요)


    ⓻ In order to reduce the data from the dimension to the () dimension, if we select the first column vectors () of and the -th leading principal submatrix () of , then becomes an matrix containing the first PC (principal components). (Look at the formula on the screen).


    .

                                                           (It's on the screen, so it doesn't have to be in the subtitles)


    그럼 는 원래 행렬보다는 size가 k×k개로 줄어듭니다. singular value들 몇 개를 활용해서 다음 분석을 시작할 지를 이 k를 적절하게 p보다 훨씬 적게 잡아서 주축 분석을 하더라도 원래 행렬 성질의 80% 또는 90%를 이해할 수 있다면, 행렬을   크기로 줄이면서도, 데이터 분석을 획기적으로 빠르게 할 수 있게 됩니다.


    Then , in comparison with the original matrix, is reduced to the size k×k. If we can understand 80% or 90% of the properties of the original matrix, even if we are doing the principal axis analysis, keeping properly k much less than p, to start the next analysis using multiple singular values, reducing the matrix to , data analysis can be much faster.


    이 를 만들어가는 과정이 PCA라고 이해하시면 되고, 그 과정의 핵심적인 idea가 SVD 라는 것이 PCA의 key idea입니다.


    We can say that the process of creating this is the PCA and that the key idea of ​​this PCA process is the <SVD>.


    [열린문제 1] 에서 주성분분석에서 특잇값분해(SVD)가 어떻게 사용되는지 정리하여 요약해 보시면 됩니다.


    [Open Problem 1] Explain how a singular value decomposition (SVD) is used in principal component analysis.


    잘 했습니다. 다음 차시에는 우리가 실제 코드를 이용해서 주성분 분석의 사례에 대해서 학습하도록 하겠습니다. 수고 많이 하셨습니다.


    Good. That's all for today. In our next lecture, we'll look at an example of principal component analysis using real 'Code'. Thank you.


    [12-2 pre]


    12주 2차시에서는 앞에서 배운 PCA 이론과 코드를 이용하여 주성분 분석의 실제사례를 학습하고, 공분산 행렬에서 시작하여, 그것의 SVD(Singular Value Decomposition)을 하여, PC score를 구하고, 주축(principal axis)를 구합니다.


    In the second session of this week, we see the real example of principal component analysis. Start from the centered data matrix, we perform its SVD (Singular Value Decomposition), compute the PC score, and find the principal axis.


    이를 통하여 PAC 주성분 분석을 하는 전체 과정을 학습한 후,  PCA와 선형회귀(linear regression) 사이의 관계를 학습하겠습니다.


    Through this, we will learn the whole process of principal component analysis, and then learn the relationship between the PCA and the linear regression.



    [12주차 2차시] 반갑습니다. 벌써 12주차 2차시입니다. PCA 학습은 오늘 마무리하려고 합니다.


    [2nd lecture, Week 12, ] Welcome to the second lecture of Week 12. Today we will complete our PСA section in Ch. 5.


    지난 1차시에 <주성분 분석>, 그 중에서도 차원축소에 대해서 학습했습니다.


    In the last lecture, we learned about theory of <Principal Component Analysis>, in particular on the dimension reduction.


    차원축소를 학습하기 위해서 Motor Trend magazine에 실린 자동차들의 데이터들을 활용했고, 그것을 분석하기 위해서, 주성분 분석이 무엇인지 소개했으며, 주성분 분석을 하기 위한 principal axes, 즉 주축을 찾는 과정, PC1, PC2, PC3 같이 principal component를 찾는 과정을 학습했습니다.


    To learn about the dimension reduction, we used <Comparison data of Passenger Cars, published in a Motor Trend magazine>, and for the analysis, we introduced what is a principal component analysis, the principal axis for the principal component analysis, i.e. the process of finding the principal axis, and with PC1 , PC2, and PC3, we studied the process of finding <Principal Components>.


    그래서 첫 번째 principal axis, 두 번째 principal axis, 즉 PC1과 PC2의 대응하는 서로 직교인 새 축을 찾는 과정을 보여드렸습니다. 그리고 이 주성분 분석을 실제 계산하기 위해서 데이터 행렬로부터 센터링하는 과정과, 센터링된 행렬을 가지고 SVD(singular value decomposition)을 해서, n×n행렬을 p×p행렬 문제로 바꿀 수 있는 이유를 바로 이 식으로 보여주었고, 즉, U의 첫 번째 p개 column 벡터들과, V의 p개 column 벡터들만 있으면, 그러면 실제 우리가 원했던 X 행렬로 표현할 수 있다는 것을 보여줬고, 여기서부터, p×p 행렬로부터 훨씬 더 작은 k차원의 행렬, k×k sub matrix를 골라서 그거에 대응하는 첫 k개의 U의 column 벡터들만 이용해서 만든 를 가지고 와 를 곱한 k×k 행렬을 만들어서 이것을 가지고 데이터 분석을 함으로써 원래 100만×100만 행렬, 200만×200만 행렬을 계산하는 대신에 적당한 size의 이 행렬을 활용함으로써, 획기적으로 계산과정을 줄이고 시간을 줄이고 비용을 줄이면서도 원래 데이터에 대한 분석과 거의 근사한 결정, 의사결정에 필요한 정보를 찾아내는 과정을 학습했습니다.


    So, I showed you the process of finding the first principal axis, the second principal axis, that is, the corresponding new axes PC1 and PC2, that are orthogonal to each other. In order to actually compute this principal component analysis for replacing the original n×n matrix with a new p×p matrix, after a centering process is performed from the data matrix, then an SVD (Singular Value Decomposition) is applied using this centered matrix. In other words, it was shown that if we only have the first p column vectors of U and p column vectors of V, then we can indeed express the desired matrix X. Hence, a much smaller k-dimensional matrix, kxk sub-matrix, is taken from the pxp matrix, and using only the corresponding first k column vectors of U, this new small matrix is created, then this new kxk matrix is ​​formed as a result of the product of and . Since analyzing the data by this process, instead of calculating the original matrix of (1 million by 1 million), (2 million by 2 million), a matrix of suitable size is used, the computation process is greatly reduced, the amount of time spent and the amount of expenses are also reduced dramatically. In addition, thanks to this, we learn how to analyze the initial data, to make approximated resonable decisions, or find some necessary informations for making a decision.


    2차시에는 <주성분 분석 사례>로 데이터는 다음 출처

    http://yatani.jp/teaching/doku.php?id=hcistats:PCA 

    에서 구했습니다. 이 데이터를 활용해서 실제로 PCA를 구현해 보도록 하겠습니다.


    In this class, <A case of principal component analysis>, data were obtained from the following source.

    [Source] http://yatani.jp/teaching/doku.php?id=hcistats:PCA

    With this open data, we will actually compute a PCA for your practice.


    실제 R 코드로 구현한 자료를 확인할 수 있게끔 제가 만들어 놨고, 그 내용을 자세히 쓰면 다음과 같습니다.


    I made it so you can check the data implemented with the real R code. More details are as follows.


      예제. 사람들이 새 컴퓨터를 선택할 때 관심을 갖는 아래 사항에 관하여 척도가 7점인 4문항의 리커트(Likert) 설문 조사를 16명에게 실시하였다. (화면의 표를 보세요)


       [Example] With a Likert scale, consisting of 4 items, each on a 7-point scale, 16 people were interviewed on the following questions. (Look at the table on the screen).


                        (1: 매우 그렇지 않다 – 7: 매우 그렇다)

    (1: strongly disagree – 7: strongly agree)


         Price        가격이 저렴하다. (The price is cheap.)

         Software     운영체제가 사용하려는 소프트웨어와 호환된다. (It is compatible with the software the operating system intends to use.)

         Aesthetics    디자인이 매력적이다. (The design is attractive.)

         Brand        유명 브랜드의 제품이다. (It is a product of a famous brand.)

                                                             <---화면에 있으니 자막에는 없어도 됨) (It's on the screen, so it doesn't have to be in the subtitles)


    16명으로부터 얻은 설문조사를 통해 얻은 (Price, Software, 디자인, Brand 에 대한 선호도) 데이터를 가지고 ‘데이터 행렬’을 만들어, 아래 Sage 코드를 이용하여 PCA를 진행합니다.  (화면의 코드를 보세요)


    Using the data (price, software, design, brand preference) obtained from a survey of 16 people, a “data matrix” is created and the PCA is conducted using the Sage code below. (Look at the code on the screen).


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    # 데이터 행렬을 입력한다.

    # Enter the data matrix.

    # 각 열은 변수 Price, Software, Aesthetics, Brand에 해당하는 데이터를 의미한다.

    # Here, each column represents data corresponding to the variables Price, Software, Aesthetics, and Brand.

    # data matrix

    X = matrix(RDF, [[6, 5, 3, 4], [7, 3, 2, 2], [6, 4, 4, 5], [5, 7, 1, 3],

                     [7, 7, 5, 5], [6, 4, 2, 3], [5, 7, 2, 1], [6, 5, 4, 4],

                     [3, 5, 6, 7], [1, 3, 7, 5], [2, 6, 6, 7], [5, 7, 7, 6],

                     [2, 4, 5, 6], [3, 5, 6, 5], [1, 6, 5, 5], [2, 3, 7, 7]])

    print("X =")

    print(X)

    print()  # 한 줄 비움 (One line indent)

    ---------------------------------------------------------------------

                                                             <-- (It's on the screen, so it doesn't have to be in the subtitles)


    데이터 행렬을 printout하고, 그 데이터 행렬을 센터링을 다음과 같이 시켜줍니다. 행의 개수, 열의 개수를 구한 다음에 센터링하는 과정을 코딩을 이렇게 해줬습니다. 그래서 센터링된 행렬 mean-centered matrix를 만들어서 printout하고, 그 행렬에 singular value decomposition을 다음과 같이 해줍니다. (화면의 코드를 보세요)


    The printed data matrix and centering process of it was shown as follows. After determining the number of rows and columns, we did code the centering process as follows. So, create a mean-centered matrix and print it out, and do the singular value decomposition over the matrix as follows. (Look at the code on the screen).


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    n = X.nrows()  # 행의 개수 (number of rows)

    p = X.ncols()  # 열의 개수 (number of columns)

    X_ctr = zero_matrix(RDF, n, p)  # 센터링된 행렬을 위해 메모리 준비 (Prepare memory for centered matrix)


    import numpy as np


    # 센터링 (centering)

    for col in range(p):

        

        v = np.array(X.column(col))

        m = mean(v)

        st = std(v)

        

        for row in range(n):

            

            X_ctr[row,col] = (X[row,col] - m)/st


    # X_ctr는 센터링된 행렬 (X_ctr is a centered matrix)

    print("X_ctr =")

    print(X_ctr.n(digits = 4))

    print()  # 한 줄 비움 (One line indent)


    # 특잇값 분해(SVD).  X_ctr = U*S*V.transpose()

    # Singular Value Decomposition(SVD). X_ctr = U*S*V.transpose()

    U, S, V = X_ctr.SVD()

    ---------------------------------------------------------------------

                                                             <---화면에 있으니 자막에는 없어도 됨) (It's on the screen, so it doesn't have to be in the subtitles)

    이 singular value decomposition을 한 다음에 principal component 들의 분산을 찾고, 그 분산의 비율을 확인한 다음에 이 분산을 시각적으로 한번 보여줍니다. Scree Plot 이 명령어를 통해서, 분산을 시각적으로 보여준 다음에, PC score를 계산합니다. 행렬, 가 ×입니다. 행렬로부터 PC score를 계산을 합니다. 첫 번째, 두 번째, 세 번째, score들을 찾고. 그 데이터를 평면에 시각화 시키는 명령어를 짜고, 파이썬 기반의 Sage 언어로 짜서 Sage로 실행을 한번 시켜보겠습니다. (화면의 코드를 보세요)


    After completing this singular value decomposition (SVD), find the variance of the principal components, check the ratio of the variance, and then visually show this variance. Use the Screen-Plot command to visually show the variance and then compute the PC score. This -matrix, that is, is computed through ×. Now, compute the PC rating by the matrix. Find the first, second, third, scores, and then write a command to visualize the data on a plane, write it in the Python-based Sage language, and run it with the Sage compiler. (Look at the code on the screen).


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    Var_PC = [(S[i, i]^2/n).n(digits = 4) for i in range(p)]  # PC의 분산 (PC variance)

    Prop_Var = [100*Var_PC[i]/sum(Var_PC) for i in range(p)]  # 분산의 비율 (Ratio of variance)

    print("PC의 분산      ", Var_PC)

    print("분산의 비율    ", Prop_Var)

    print()  # 한 줄 비움 (One line indent)


    # Scree Plot. 분산을 시각적으로 보여줌. 몇 차원으로 줄일지 결정.

    # Scree Plot. Visually show variance. Decide to what dimension to reduce.

    show(list_plot(Var_PC, plotjoined = True, color = 'red', axes_labels = ['', 'Variances']))

    print()  # 한 줄 비움 (One line indent)


    # PC score를 계산한다.  Z = U*S

    # Calculate the PC score.  Z = U*S

    PC_score = U*S

    PC12 = PC_score[:, 0:2]  # 시각화를 위해 첫 번째, 두 번째 성분까지만 사용 (Use only the first and second components for visualization)


    # 데이터를 평면에 시각화 한다. (Visualize the data on a plane.)

    point([PC12.row(i) for i in range(n)], color = 'red')

    ---------------------------------------------------------------------

            <---화면에 있으니 자막에는 없어도 됨) (It's on the screen, so it doesn't have to be in the subtitles)


    이제 결과를 살펴보겠습니다.

    Now let's look at the results.


    그 데이터 행렬들을 먼저 printout하고, mean-centered 행렬을 확인한 후에, principal component의 분산을 확인하고, 분산의 비율을 한번 본 후에 Scree Plot을 통해서 분산이 어떻게 변해가는 지를 보면 principal component들의 영향을 확인할 수 있고, 그것을 스크린으로 찍어주고, 그리고, 이 분산 확률을 가만히 보시면, 여기서 보시다시피 principal component가, 첫 번째 principal component, 두 번째, 세 번째, 네 번째 component가 있는데, principal component 중에서 첫 번째와 두 번째가 60%와 24%이니까 거의 85%, 85%의 분산이 첫 번째와 두 번째 component, PC1와 PC2에 영향을 받는다는 것을 이해할 수 있습니다. 그림으로 봐도 첫 번째 두 번째를 거치면 벌써 이만큼 까지 왔습니다. (화면의 수식을 보세요)


    First, print the data matrix, check the mean-centered matrix and variance of the principal component, then using Scree-Plot we can see how the variance changes and we can also find out the effect of the principal component. If we take a careful look on this variance probability, we can see here four principal components, among them the first and the second components are 85%, which means that almost 85% of the variance is influenced by this first and second components, namely PC1 and PC2. Even if we look at the .\PICture, we will see that we have already achieved such an impressive result after using the first and second components. (Look at the formula on the screen)


    X =

    [6.0 5.0 3.0 4.0]

    [7.0 3.0 2.0 2.0]

    [6.0 4.0 4.0 5.0]

    [5.0 7.0 1.0 3.0]

    [7.0 7.0 5.0 5.0]

    [6.0 4.0 2.0 3.0]

    [5.0 7.0 2.0 1.0]

    [6.0 5.0 4.0 4.0]

    [3.0 5.0 6.0 7.0]

    [1.0 3.0 7.0 5.0]

    [2.0 6.0 6.0 7.0]

    [5.0 7.0 7.0 6.0]

    [2.0 4.0 5.0 6.0]

    [3.0 5.0 6.0 5.0]

    [1.0 6.0 5.0 5.0]

    [2.0 3.0 7.0 7.0]


    X_ctr =

    [  0.8485 -0.04218  -0.7500  -0.3866]

    [   1.317   -1.392   -1.250   -1.511]

    [  0.8485  -0.7170  -0.2500   0.1757]

    [  0.3804    1.307   -1.750  -0.9489]

    [   1.317    1.307   0.2500   0.1757]

    [  0.8485  -0.7170   -1.250  -0.9489]

    [  0.3804    1.307   -1.250   -2.074]

    [  0.8485 -0.04218  -0.2500  -0.3866]

    [ -0.5559 -0.04218   0.7500    1.300]

    [  -1.492   -1.392    1.250   0.1757]

    [  -1.024   0.6327   0.7500    1.300]

    [  0.3804    1.307    1.250   0.7380]

    [  -1.024  -0.7170   0.2500   0.7380]

    [ -0.5559 -0.04218   0.7500   0.1757]

    [  -1.492   0.6327   0.2500   0.1757]

    [  -1.024   -1.392    1.250    1.300]


    PC의 분산       [2.278, 0.9011, 0.4356, 0.1348]

    분산의 비율     [60.76, 24.03, 11.62, 3.596]


                                                             <---화면에 있으니 자막에는 없어도 됨) (It's on the screen, so it doesn't have to be in the subtitles)


    # 따라서 PC1, PC2가 원 데이터 분산의 85% 정도를 보존한다고 이해할 수 있다. ■


    # Hence, from here it can be understood that PC1 and PC2 retain about 85% of the original data variance. ■


    이런 경우에는 데이터가 4차원이었지만, 2차원으로 줄여서 분석을 하더라도, 거의 85% 정도 데이터의 성질들에 대한 분석이 이루어지므로 큰 그림에서는, 의사 결정을 하는데 도움이 된다는 것입니다. 그렇다 그러면 시간을 절반, 1/4 이상으로 줄일 수 있다는 뜻입니다.


    In this case, the original data was four-dimensional, but even if we reduced it to two-dimensional and analyzed, it also retains about 85% of the properties of the original data. So, in general, they are useful for a decision-making. This means that we can cut the time in half, or at least more than a quarter.


    이것이 4차원을 2차원으로 줄였으니까 이 정도지만, 예를 들어 100만 차원을 1000차원으로 줄였다 그러면 이 비용이 절감되는 액수는 상상할 수 없을 것이고, 그 계산에 소요되는 시간이 줄어드는 것도 상상할 수 없게 줄어들 것입니다.


    We got such result, since we reduced the 4th dimension to the 2nd dimension. However, for example, if we reduce the 1 millionth dimension data to a 1000 dimension, then we will significantly reduce the amount of economic costs and time.


    이런 이론을 활용하기 때문에 인공지능에서 우리가 명령을 해서 ‘이것 좀 알아봐달라’ 그리고 파악을 하면 바로 그 답을 얻을 수 있는 것입니다. 실제 100만×100만 데이터 행렬을 처리한다고 그러면 그렇게 빠른 시간에 답을 들을 수가 없습니다.


    However, if we decide to process with a 1 million x 1 million data matrix, then we will not be able to get the result in such a short time. Through the use of this theory, by giving the artificial intelligence the command "Please check this", we can immediately obtain and analyze the result.


    이어서, 아래 주성분의 고윳값에 대해선 line plot을 활용하여 데이터의 차원을 줄일 때, 주성분을 몇 개로 선택할지 결정하는 과정을 볼 수 있습니다. 고윳값을 가장 큰 값에서부터 크기 순서대로 정렬하여, 대개 기울기가 꺾이는 부분의 “elbow point” 왼쪽에 있는 성분까지를 선택합니다. 따라서 15% 정도의 정보는 잃어버리더라도, PC1, PC2만 택하여 2차원 데이터로 차원을 줄 일 수가 있습니다. 이를 활용하여 데이터를 시각화하거나 다른 데이터 분석기법을 적용할 수 있는 것입니다. 여기서 보면 한번, 두 번 꺾이면, elbow point, 여기 두 축 정도만 되면 전체적인 분포의 85%를 이해할 수 있다는 의미입니다.  (화면의 표를 보세요)


    Next, for the eigenvalue of the principal components below, we can see the process of deciding how many principal components to select when reducing the dimension of the data using a line plot. Eigenvalues are sorted in the order of magnitude from the largest component to the component to the left of the “elbow point”, usually where the slope is bent. Therefore, even if we lose about 15% of the information, we can reduce the dimension to 2D data by selecting only PC1 and PC2. We can use this to visualize your data or to apply other data analysis methods. If we look at these figures, we can see that if there are one or two bends, that is, an elbow points, it is enough to have two axes to understand about 85% of the overall distribution. (Look at the table on the screen).


                                         묶음 개체입니다.

                                                                             <---화면에 있으니 자막에는 없어도 됨)



    연습문제로, 본인 전공에서 찾은 데이터행렬을 가지고, 위의 코드를 활용하여 위의 주성분분석(PCA)의 알고리즘을 그 행렬에 적용해보시고 한번 문의 게시판에 토론해 보시면 이해에 도움이 될 것입니다.


    As an exercise, take a data matrix from sources related to your major, then use the code above to apply a Principal Component Analysis (PCA) algorithm to this matrix, and then, post the results on the QnA Board for discussion. This will help you to understand the to.\PIC.


    마지막으로 <주성분 분석과 공분산 행렬>절에는 별 표시 *(star)가 붙어 있습니다. 이는 필수가 아니라 선택적으로 학습하는 절이라는 의미입니다.


    Finally, let's look at the optional <Principal Component Analysis and Covariance Matrix> section, which is marked with an * (asterisk). This means that this section is provided for your optional study.


    주성분 분석은 원 데이터의 분포를 최대한 보존하면서 고차원 공간의 데이터들을 저차원 공간으로 변환하는 기법입니다. 그런데 원 데이터의 분포에 대한 정보는 공분산 행렬에 담겨 있습니다. 그래서 주성분 분석과 공분산 행렬과는 아주 밀접한 관계가 있습니다. 여러분들이 선형대수학을 배운 이유가 여기 있습니다. 주성분 분석을 하려면 이 행렬에 SVD를 할 수 있어야 되는 것입니다.


    Principal component analysis is a technique that converts data from high-dimensional space to low-dimensional space while preserving the distribution of the original data as much as possible. However, the information about the distribution of the raw (original) data is contained in the covariance matrix. So, there is a very close relationship between principal components analysis and covariance matrices. This is one of the main reasons why we learned Linear Algebra. In order to do principal component analysis, we need to be able to apply SVD on this matrix.


      데이터 행렬 를 센터링한 행렬을 가지고,  공분산 행렬을 만들고,  공분산 행렬이 대칭행렬이기 때문에 대각화가 가능하고, 고윳값은 언제나 실수입니다.


    Center the data matrix X, then create a covariance matrix based on this matrix, and since the covariance matrix is a symmetric matrix, it is also diagonalizable, which means that its eigenvalues are always real numbers.


    주성분 분석의 목표는 ‘공분산 행렬로부터 얻은 정보’를 최대한 보존하는 ‘더 적은 개수의 새로운 변수들’을 찾으려는 것입니다. 이를 이용하여 훨씬 더 작은 크기의 행렬로 rank를 reduction 하는 것이 우리의 목표입니다.


    The goal of principal component analysis is to "Find the smallest number of new variables" that can "preserve the information from the covariance matrix as much as possible". Our goal is to use this method to reduce the rank to (MAKE) a much smaller matrix.


    이 행렬로부터 다음 관계를 얻습니다. 그러면 SVD에서 사용하던 직교행렬을 가지고 공분산 행렬을 표현할 수 있게 되는 것입니다.


    From this matrix we get the following relation. Thus, the covariance matrix can be expressed with the orthogonal matrix used in SVD.


    이때, 를 양변에 곱해주면, 직교행렬이니까, 이렇게 되고, 잘 보면, 다음과 같은 관계가 성립되는데, 이게 의미하는 바가 무엇이냐면, Ax=x 같이 고윳값과 고유벡터의 관계입니다. 즉, 는 의 고윳값, 는 고유벡터가 되는 것입니다. 는 번째 PC의 분산을 나타내며, 는 번째 주축을 나타냅니다.


    Now, if we multiply on both sides by , since this is an orthogonal matrix, we get the following . If you look closely this relationship, you will see that it is like Ax=x, which means the eigenvalue- eigenvector relation. That is, is the eigenvalue of , and is the eigenvector. represents the variance of the - th PC, and represents the -th principal axis.


    고윳값, 고유벡터에 관하여 여기서는 간단히 살펴보았지만, 여러분들이 앞에서 배운 singular value와 singular vector가 바로 이 경우에, 대칭행렬인 경우에는 고윳값과 고유벡터와 같은 개념이라고 이해하시면 됩니다.


    So far, we have briefly considered the concept of eigenvalue ​​and eigenvector, later we will study the concept of singular value and the singular vector, however, in our case, that is, in the case of a symmetric matrix, these concepts are identical.


    마지막으로, 공분산 행렬에 대한 주성분분석을 하여, 차원을 축소하는 과정에 대해서 여러분들이 이해한 대로, 데이터 행렬을 mean-centered 행렬로 만들고 그 다음에 singular value decomposition을 이용해서 이 행렬을 만든 다음에 principal component를 골라서, 그래서 이제 dimension을 n×n에서 p에서 그 다음에 k까지로 줄여서 dimension reduction 과정을 이해하셨으면 충분합니다.


    Finally, perform principal component analysis on the covariance matrix. Next, follow the dimension reduction process that we learned earlier. Convert the data matrix to a mean-centered matrix, after that, use the singular value decomposition, and then select the principal component from the resulting matrix, as a result we will reduce the dimension from nxn to pxp, and then to kxk. This is enough to understand the concept of the dimension reduction.


    그 다음에 이 주성분 분석이 통계학의 선형회귀(linear regression)와 어떤 관계가 있는지를 star 절에서 제공하였습니다. 이것에 관해 추가설명을 하면, 선형회귀에서는 이런 데이터가 일차함수에 의해서 계산된 점들 사이의 오차가 생깁니다. 오차를 최소화 하는 방향으로 계수 , 를 결정했는데, 주성분 분석에서는 각 데이터에서 일차함수 에 이르는 거리가 최소가 되도록 와 를 결정한 것입니다.


    Then, in the section marked with an asterisk, we provided information on how this principal component analysis relates to statistical linear regression. This is also explained by the fact that since linear regression uses a linear function for this data, an error occurs between the computed points. In a linear regression, the coefficients and were determined in the direction of minimizing the error, however, in the analysis of the principal components and are determined in such a way that the distance from each data point to the linear function   is minimal.


    마찬가지로 그 때, 주성분 분석을 사용하여 얻어진 직선이 더 데이터에 가까이 있는 것을 선형회귀(linear regression) 보다도 더 효과적인 것을 볼 수 있습니다.


    Also, here we may notice that the straight line obtained by a principal component analysis is located closer to the data than the line obtained using the linear regression, hence we can say that the principal component analysis method is more efficient than the linear regression method.


     아래 그림은 두 가지 선형회귀를 이용한 것과 주성분 분석을 이용해 얻어진 직선을 같이 보면, 주성분 분석이 앞에서 배웠던 선형회귀보다도 더 효과적인 것을 확인할 수 있습니다. 즉 최소제곱직선과 회귀분석에서 사용하는 선형회귀 사이에는 유사한 점도 있고 차이점도 있습니다. 그것을 오늘 배운 PCA 와 비교해보는 문제를 생각해보시고 이 절을 마무리 하시면 되겠습니다.


    In the figure below, two straight lines are shown together, the first one obtained using a linear regression and the second one using a principal component analysis. If you look closely at the figures, you will notice that the principal component analysis method is more effective than the linear regression method that we studied earlier. In other words, there are similarities and differences between the least squares straight line and the linear regression used in regression analysis. Consider the comparision of these methods to the PCA we studied today. At this point we can conclude this section on PCA.


    이제, 이번 학기 14주 강의 중에서 가장 중요한 내용인 PCA에 대해서 학습했습니다. 아이디어는 고차원 데이터는 계산도 어렵고 시각화도 어려워서 분석하기가 힘들고, 시간도 오래 걸리니까, 이 원 데이터의 분포를 가능한 최대한 유지하면서 데이터의 차원을 우리가 원하는 만큼 우리가 다룰 수 있는 만큼 우리가 원하는 시간 안에 답을 계산해서 찾아서 우리에게 제공해줄 수 있을 만큼 차원축소(dimension reduction)이 필요하다는 것입니다. 이를 위해서 PCA를 사용하고, PCA하는 것은 공분산 행렬을 만들어서 SVD를 해서 얻을 수 있다는 내용을 학습했습니다.


    Now, we have finished studying one of the most important sections of this course, which is PCA. The idea is that a multi-dimensional data is difficult to compute and visualize, hence difficult to analyze and also takes a long time to compute. Therefore, we need a dimension reduction method that can, in a short period of time, provides us with data in an acceptable dimension for us, while preserving the distribution of the original data as much as possible. Such a method is the PCA that we studied today. As we know, in order to use PCA, we need to create a covariance matrix, and then apply SVD on it.


    그 과정을 설명하기 위해서 1절에서는 차원축소, 2절에서는 주성분 분석, 계산, 사례를 보여주었고, 이제 5절에서는 주성분 분석과 공분산 행렬, 마지막에 주성분 분석과 앞에서 배웠던 선형회귀와의 관계를 설명했습니다.


    To explain this process, in the first lecture of this week we considered at the concept of “Dimension reduction”, in second lecture we studied Principal component analysis and looked at a real example of calculation. In this section 5, we talked about the relationship between the principal component analysis and the covariance matrix, and finally, we explained the relationship between principal component analysis and the previously studied linear regression.


    그 내용을 보면, 차원 축소, 이런 데이터로부터 주축을 찾는 과정을 설명했고 SVD 해서 그것이 차원축소와 어떻게 관계가 있는지, 차원축소를 했을 때, 효과가 4개의 축을 2개의 축으로 줄여서 계산하더라도 즉, 두 축만 사용해서 새로운 축에 대한 데이터(data), 즉 관련된 행렬에 대한 분석만 하더라도 85% 성질을 보존하는 분석이 가능하다는 것을 확인한, 수학적으로 확인한 내용입니다.


    Briefly summarizing all of the above, we explained dimension reduction, and the process of finding the principal axis for a given data, and how it relates to downsizing using SVD. It is mathematically confirmed that even if the obtained effect when we compute a dimension reduction, is a reduction from 4-axes to 2-axes, that is, using only two axes, analyzing the data with a new axis, that is, only by analyzing the matrix associated with it, it is already possible to carry out an analysis with about 85% preservation of the original data.


    이제 다음 시간에는 인공 신경망(Artificial Neural Network)에 대해서 학습하도록 하겠습니다. 수고하셨습니다.


    Next time, we will learn about Artificial Neural Networks (ANN). Thank you.

     

     

    Week 13. Introductory Mathematics for Artificial Intelligence

     

                인공지능을 위한 기초수학 입문


    13. 인공신경망 (Artificial Neural Network) 145

    13.1 신경망

    13.2 신경망의 작동원리

    13.3 신경망의 학습

    13.4 오차 역전파법


    [13-1 pre]


    여러분 반갑습니다.


    Welcome!! 


    지난 12주 동안 우리는 인공지능을 학습하는데 필요한 행렬, 미적분, 통계에 대한 기초 지식과 PCA를 학습하였습니다.


    Over the past 12 weeks, we have learned the basic knowledge of matrices, calculus and statistics and PCA thar are essential to understand AI.


    이번 주 수업에서는 지금까지 배운 지식을 활용하여 인공신경망(Artificial Neural Network)이 어떻게 작동되는지에 대해서 학습하도록 하겠습니다.


    In this lecture, we will use the knowledge we have learned so far to learn how the Artificial Neural Network works.


    신경망(neural network)은 신경계의 기본단위인 뉴런을 모델화 한 것으로, 이미지 인식 등의 업무를 잘 수행합니다. 이번 시간에는 특히 앞에서 배웠던 행렬의 연산, 행렬의 분해, 특히 SVD와 경사하강법을 활용하여 <인공신경망의 작동원리>를 설명하도록 하겠습니다.


    A neural network is a model of neurons, the basic unit of the nervous system, and performs tasks such as image recognition. In this lecture, we will explain how the artificial neural network works. It will done by using the matrix operation and the gradient descent, which we learned earlier.

     

    순서는 1절에서는 신경망, 2절에서는 신경망의 작동 원리, 3절에서는 기계학습, 4절에서는 오차 역전파법(back propagation)에 대해서 학습하겠습니다.


    We will learn about neural networks in section 1. In section 2, we will explain how the neural networks works. Section 3 is for the learning of neural networks. And the backpropagation will be covered in section 4.


    우리의 뇌 속의 신경망은 그림과 같이 시냅스, 수상돌기, 축삭돌기 등을 거치면서 작동합니다. 이것을 다음 그림과 같이 input이 있고 weight가 주어지고, 그 다음에 층(layer)들이 주어진 후에, 그 과정을 통해서 output으로 나타나는 과정으로 인공신경망을 디자인한 것입니다.


    The neural network in our brain works through synapses, dendrites, and axons as shown in the .\PICture. And this is an artificial neural network designed as an output through the process after computation with input data and weights, and then layers are given as shown in the following figure.


    그래서 input으로 데이터들 가 들어가면 output으로 ‘이 사람은 몸무게가 어느 정도 되어야 하고, 혈압은 어느 정도 되는 게 정상이다’라는 output이 나오게 됩니다.


    When a data is entered as an input, an output stating ‘this person should have a certain weight and a certain blood pressure is normal.’ is obtained.


    이는 벡터가 들어가서 벡터가 나오는 것이니까, (주어진 벡터에) 어떤 행렬을 곱해서 얻어진 결과 (벡터)로 이해할 수 있습니다. 우리가 그 행렬, 즉 함수의 모양이 무엇인지를 찾아내는 것이 key idea입니다.


    Since this is a vector-in and vector-out, so it can be understood as the resulting vector was obtained by multiplying a matrix to a given vector.


    The key idea is for us to find out what the shape of the function which is a matrix.


    또, 신경은 일정한 자극이 주어졌을 때 처음에 작은 자극이 주어졌을 때는 반응하지 않지만, 그 강도가 점점 세져서 어느 정도를 넘어서면 갑자기 ‘아야!’하고 반응하듯이 그에 대응하는 함수가 Heaviside function인데, Heaviside function은 미분가능하지 않기 때문에 우리가 지금까지 배운 수학적 지식을 효과적으로 활용할 수 없어서 그에 가장 근사한 활성화(activation) 함수로 Sigmoid 함수나 ReLU 함수 등을 활용합니다.


    In general, the nerve does not respond when a small stimulus is initially given. But the intensity of the nerve is gradually increased, and when it exceeds a certain level, it suddenly responds with 'Ouch!'. The corresponding function is a Heaviside function. Since the Heaviside function is not differentiable, we cannot effectively utilize the mathematical knowledge we have learned so far, so we use the Sigmoid function or ReLU function for the activation function.


    이에 필요한 기초지식을 학습한 후에, 역 전파법, 즉 함수 box안의 층(layer)들 사이에서 일어나는 일을 설명합니다. 이 층(layer)들이 여러 개 있으면, Deep learning이라고 합니다. 전체 기계학습(Machine learning) 중에서 Deep learning이 작동하게 하는 은닉층(Hidden layer)들과 은닉층들 사이에서 어떻게 함수를 생성하고 update해 가는지, 그 과정을 <오차 역전파법(back propagation)>이라고 합니다. 이  오차 역전파법의 수학적 원리를 자세하고도 쉽게 설명 할 것입니다.


    When there are many hidden layers in AI, it is called a deep learning. Among all machine learning techniques, the process of creating and updating weights between hidden layers is called a <Backpropagation>. The mathematical principle of this <Backpropagation> will be explained in detail and easily. After learning the basic knowledge required, we explain the <Backpropagation>, that explains what happens between the layers in the neural network.


    [13주차 1강]


    반갑습니다. 일반인을 위한 K-MOOC [인공지능을 위한 기초수학 입문] 13주차, <인공신경망(Artificial Neural Network)>에 대해서 학습하겠습니다.


     [Lecture 1, Week 13] Welcome to the first lecture of the 13th week lecture, K-MOOC [Introductory Mathematics for AI]. Today, we will study <Artificial Neural Network>


    1절은 신경망입니다. 사람의 뇌에 있는 뉴런(신경세포, neuron)은 혈액 중의 아미노산으로부터 신경전달물질을 만들어 냅니다. 신경의 자극은 한 방향으로 일어나며, 만들어진 신경전달물질을 이용하여 주위의 다른 신경세포들에 신호를 전달합니다. 하나의 신경세포는 여러 개의 다른 신경세포들로부터 전달받은 신호(입력신호)에 반응하여 세포막이 어떤 임계값(역치, threshold value) 전위에 도달하면 다음 신경세포에도 일정한 크기의 신호(출력신호)를 전달하게 됩니다.


    <Section 1> deals with the <Neural network>. Neurons(nerve cells) in a human brain make <Neurotransmitters from amino acids in the blood>. Nerve stimulation occurs in one direction, and the signals are transmitted to other nerve cells around it by using the generated neurotransmitter. One nerve cell responds to signals (input signal) received from several other nerve cells, and if the cell membrane reaches a certain threshold potential, it transmits a signal (output signal) of a certain size to the next nerve cell.


    인간의 뇌는 대략 1,000억 개의 뉴런을 가지고 있으며, 이를 통하여 어려운 이미지 인식과 같은 업무를 아주 잘 수행합니다. 이 그림이 뇌 속에 있는 신경망의 모습니다.

     

    The human brain has about 100 billion neurons and performs difficult tasks very well, such as image and pattern recognition. This is a .\PICture of a neural network in the brain.


    이것을 모델링 한 것이 바로 인공 신경망(Artificial Neural Network)입니다. 여러 가지 input이 들어가서 어느 정도 자극의 레벨(level)을 넘어서면 반응을 일으키는 시스템입니다.


    The mathematical model of this neural network is <the Artificial Neural Network>. It is a system that causes a response when various inputs exceed a certain level of stimulation.


      이 신경망(neural network)은 신경계의 기본 단위인 뉴런을 모델화한 것인데, 하나의 인공 뉴런(노드, node)에서는 다수의 입력 신호 를 받아서 하나의 신호를 출력합니다. 이는 실제 뉴런에서 전기신호를 내보내 정보를 전달하는 것과 비슷합니다. 이때 뉴런의 돌기가 신호를 전달하는 역할을 하듯이 인공 뉴런에서는 가중치(weight) 가 그 역할을 합니다. 각 입력신호에는 고유한 가중치가 부여되며 가중치가 클수록 해당 신호가 중요하다고 볼 수 있습니다. 이 그림이 인공신경망 한 개의 예입니다.


    This neural network is a model of a neuron, which is a basic unit of the nervous system. One artificial neuron(node) receives multiple input signals and one output signal. This is similar to transmitting information by sending electrical signals in real neurons. At this time, just as the axon of a neuron transmits signals, in an artificial neuron, the weight plays that role. Each input signal has its own weight, and a signal with larger weight will be a more important signal. This figure is an example of an artificial neural network.


    이제 신경망이 작동하는 원리에 대해서 학습하겠습니다. 아래와 같이 하나의 노드는 입력 신호를 받아 결과를 전달해주는 함수로 볼 수 있습니다. 따라서 입력신호 를 받아 를 출력하는 함수 로 나타낼 수 있습니다. 여기서 는 약간의 편향(bias)을 의미합니다. 예를 들어, 입력신호 를 받으면, 미리 부여된 가중치 와 계산 후 그 결괏값이 임계값 를 넘으면 1을 출력하는 모델입니다.


    Now we will study how this artificial neural network works. As shown below, one node can be seen as a function that receives an input signal and transmits the result. Therefore, it can be expressed as a function that receives an input signal and outputs . Here, is a bias. For example, it is a model that outputs 1, when the result of computation with an input signal and the weights assigned in advance exceed the threshold value.


                                  output 1  ()

      

    이 때, 그 과정을 함수로 이해하면, 가 입력이 되고 가 출력이 되는 과정입니다. 가장 간단한 모델이 모양입니다.


    At this moment, if we consider this process as one function, it is the process where is an input and is an output. The simplest model is .


      이것을 수식으로 표현하면 에 라는 가중치를 주고 1에 라는 가중치를 주어서 가 로 나타나는 셈입니다.  와 이 들어가면 가 로 출력(output)되는 모델입니다.


    This can be expressed as an equation, , by giving a weight to x and a weight to 1. It is a model where the output is y when and 1 are inputs.


      같은 방법으로 입력신호가 두 개인 경우에는, 과  를 입력했을 때, , 또 입력신호가 세 개인 경우에는 를 입력했을 때, 선형함수 (linear하게) 로 표현할 수 있습니다.


    In the same way, it can be expressed as a linear function when there are two input signals, , and the output is when we have three input signals, .


      그림으로 보면, 과  가 입력되면 가중치 가 주어져서 output이 로 나타나고, 입력신호가 세 개인 경우에는 출력이 이란 output 으로 나타나는 것을 볼 수 있습니다.


    In the figure, when and are inputs with given weights , and the output is . And when there are three input signals, and , we may have the output .

     

      만일에 출력이 하나가 아니라 두 개인 경우는 다음과 같이 행렬의 곱을 이용하여 나타낼 수 있습니다. 편의상 bias 편향을 0이라고 두면, 웨이트(weight) 가 여러 개가 되니까, 라고 놓겠습니다. 를 입력층의 번째 노드에서 출력층의 번째 노드로 연결된 가중치라고 합니다. 그러면 에서 으로 가는 weight, 그리고 에서 로 갈 때의 weight를 , 라고 두고, 에서 로 갈 때의 weight를 라고 놓는다면, 가 입력되면 각각에 가중치를 반영하여 라는 벡터로 나타나는 모습을 모델링 할 수 있습니다.  (화면의 수식을 보세요)


    If there are two outputs, not one, it can be expressed using the product of matrix as follows. For convenience, if the bias is 0 for simplicity, since there are multiple weights, We may put weights as 's. This is called the weight connected from the th node in the input layer to the th node in the output layer. Then, if we put the weight going from to and the weight going from to are and . And the weight going from to is . We can model these as vectors and can be found as the output when we multiply this weight matrix and to the input vector ( ).

                               


                                                             <---화면에 있으니 자막에는 없어도 됨)


    즉, 가 들어가면 4×1 벡터를 입력하면 거기에 함수를 거쳐서 라는 벡터로 나오는 선형연립방정식을 생각할 수 있습니다. 그러면 자연스럽게 weight들을 집어넣어 은 로 잡아줄 수 있고 같은 방식으로 두 번째 도 정의하는 식으로 신경뉴런을 행렬을 이용해서 표현할 수 있습니다.


    In other words, if you input (), vector, we can think of a system of linear equations that gives an output vector () through a function. Then, we can put the weights and express = and, we can express in a similar way by using the weight matrix.


      이것을 보통 통계나 인공지능에서는 표현을 보통 4차원 벡터가 들어가면 1×4 벡터가 들어가면 4×2 행렬을 곱해서 (, )라는 1×2 벡터나 나온다는 식으로,  전치행렬(transpose)을 취한 모습으로 많이 사용합니다. (화면의 수식을 보세요)


    In general, in statistics or artificial intelligence, the expression, , is usually expressed as a transpose matrix such as,   as vector is the input, a matrix is multiplied it to yield a vector.

                    


                                                             <---화면에 있으니 자막에는 없어도 됨)


    이 방법이 어떤 의미에서는 더 직관적이기도 해서, 인공지능과 통계 분야의 현실적인 응용에서는 이런 표현을 사용한다고 볼 수 있습니다. 이런 식으로 우리가 뉴런을 함수, 특히 선형함수를 이용해서 모델링을 할 수 있다는 것을 이해할 수 있습니다.


    This method is more intuitive in some sense, so practical applications can be seen the in the field of artificial intelligence and statistics. In this way, we can understand that we can model neurons using functions, especially linear functions (Matrix).


    뉴런에서 임계값 이상의 자극이 주어져야 감각이 반응하는 것처럼, 바늘로 찌를 때, 처음에 살살할 때는 약간 느껴지지만 별로 반응하지 않다가 깊이 들어가서 신경을 건드릴 때, 신경에 가까워지면서 ‘아야!’하고 이렇게 반응을 보이듯이, 인공 뉴런에서도 다수의 입력신호가 주어지면, 미리 부여된 가중치와 계산을 한 후 그 총합이 정해진 임계값을 넘으면, 1을 출력하고 넘지 못하면 0 (또는 -1)을 출력하게 됩니다. 이때 출력을 결정하는 함수를 활성화 함수(activation function)라 하는데, 활성화 함수의 대표적인 예로 Sigmoid 함수 가 있습니다. 그 Sigmoid 함수는 ‘자극이 주어져도 반응하지 않다가 어느 순간을 넘으면 ‘아야!’하고 반응하게 되는 미분가능하지 않은 Heaviside function 대신, Heaviside function과 거의 같은 behavior을 갖고 있지만 연속이면서 동시에 미분가능한 그리고 도함수가 특히 특별히 편리한 성질을 가지고 있습니다.


    Just as the sensation responds only when a stimulus is above the threshold in a neuron. For example when we are pricked with a needle, it is felt a little and we don't respond very much when it is touched at first, then it goes deep and gets closer to the nerve, we react like ‘Ouch!’ In an artificial neuron, if a number of input signals are given, calculate with the weights assigned in advance, and if the sum exceeds a predetermined threshold, outputs 1, if not, outputs 0 (or –1). At this time, the function that determines the output is called <the Activation Function>. <A role model of an activation function is the Sigmoid function, >. The Sigmoid function has almost the same behavior as the Heaviside function, but it is continuous and continuously differentiable. Especially, its derivative has particularly convenient properties.


    (그래프를 보면) 빨간색이 Heaviside function이고 파란색이 Sigmoid function입니다. 자극이 반응하지 않다가 어느 순간에 ‘아야!’하면서 반응하는 것을 표현한 것입니다.

     


    (See the graph) This Red graph is an Heaviside function and this Blue graph is a Sigmoid function. These graphs show that a stimulus does not respond for a while and reacts at certain point by saying ‘Ouch!’


    그것을 코드를 이용해서 실제로 한번 그려보겠습니다. 앞에 Sigmoid function이 파란색으로 그려지고 Heaviside function이 빨간색으로 그려지는 것으로 볼 수 있습니다. 여기서는 이 함수를 activation function로 사용합니다.


    Let’s draw it using a <Code>. A Sigmoid function is drawn in blue and an Heaviside function is drawn in red. Here, we use this Sigmoid function as the activation function in our lecture.


      Sigmoid 함수 는 다음 성질을 갖습니다. 첫째, 는 모든 에 대해서 0과 1사이에 있고, 일 때는 이고 일 때는 값을 갖습니다.

    The Sigmoid function has the following properties. First, is between 0 and 1 for all , and if , and if , .


      그리고 Sigmoid 함수의 미분 공식이 중요합니다. 의 도함수(derivative)를 구하면 연속이고 미분가능하며 derivative가 존재하는 그런 미분가능한(nice, smooth) 함수이기 때문에, 도함수를 계산해 보면 간단한 공식에 의해서 인 것을 알 수 있습니다. 그러면 를 구할 때에도 를 갖고 구할 수 있고, 를 구할 때도 마찬가지로 만 가지고 구할 수 있으니까, 도함수를 구하는 문제가 실제 미분을 전혀 하지 않고도 를 대입해서 얻을 수가 있게 되는 것입니다.


    And this differential formula of the Sigmoid function is important. If we compute the derivative, we can find out . Then, we can get from easily, also we can get with only , so the problem of finding the derivative can be done easily by substituting without doing any derivative at all.


      실제 이 식이 맞는지 확인해봅시다. 를 Sigmoid 함수로 정의해주고 를 구해보고, 또 가 와 같은지 확인해보면 가 이 식과 같이 나오고, 그리고 둘이 같다는 것, 즉 좌변과 우변이 같다는 것도 바로 확인할 수 있습니다.


    Let’s check whether this is actually correct or not by using a 'Code'. If we define as the Sigmoid function and find out and check that it is equal to , then comes out like this, and we can immediately confirm that the two are the same, that is, the left and right sides are the same.


      이 Sigmoid 함수는 Heaviside function(계단함수, step function)와 비교할 때, 출력신호를 극단적인 값 (0 또는 1)이 아니라 연속적인 0과 1사이의 값으로 정규화 하여 전달해줍니다. 또한 Heaviside function(계단함수, step function)은 일 때 미분가능하지 않으나 Sigmoid 함수는 모든 실수 에서 미분이 가능하며, 계산도 아주 간단합니다. 이외에도 활성화 함수로 ReLU(Rectified Linear Unit) 함수, 하이퍼볼릭 탄젠트() 등이 사용됩니다. 우리는 우선 여기서는 Sigmoid 함수에 집중해서 이론을 전개하도록 하겠습니다.


    Compared with the Heaviside function(step function), this Sigmoid function normalizes the output signal to a continuous value between 0 an 1, rather than an extreme value (0 or 1). In addition, the Heaviside function is not differentiable when , but the Sigmoid function is differentiable on all real numbers and the computations are very simple. Besides, ReLU(Rectified Linear Unit) function, hyperbolic tangent(), etc are more used as activation functions in real life problems. For now, we will focus on the Sigmoid function and develop the theory.


      이제 이 과정을 정리해보겠습니다. 먼저 인공 뉴런은 입력 데이터 ()를 다른 인공 뉴런으로부터 혹은 외부로부터 입력받습니다. 입력이 들어오면 인공 뉴런은 입력받은 데이터와 가중치를 결합하여 아까 보았듯이 output을 (하나의 값을 가중치와 일차결합을 해서 거기에 bias까지 보태서) 만듭니다.


    Let’s organize this process. First, artificial neurons receive input data() from other artificial neurons or from outside. When an input comes in, the artificial neuron combines the input data with the given weights and creates an output (by linearly combining a single value with the weight and adding the bias).


       ⓶에서 구한 값을 활성화 함수에 대입하면 출력값(output)이 나오게 됩니다. 어느 정도 값을 넘어서면 반응이 되고, 작은 자극이 들어갔을 때는 반응을 보이지 않는 그런 모습을 갖게 될 것입니다.


    Substituting the value obtained in ⓶ into the activation function, output is returned. When it exceeds a certain value, there is a response, and when a small stimulus was given (input), it will have an output that does not respond.


      예를 들어, 3개의 입력데이터 ()와 가중치가 주어진 인공신경세포를 활성화 함수 시그마 또는 , 위에서는 라고 썼습니다. 그림에서는 시그마라고 쓴 이것이 활성화 함수입니다.


    For example, an artificial neuron with 3 input data (), and the weights is written as the activation function sigma or . In the figure, it is written as sigma and this is the activation function.


      이 그림을 도식화하여 보면, 아래 그림과 같고, 이 과정을 통해서 출력값을 얻게 되는 것입니다. 즉, ()가 들어가면서 가중치 를 곱해서 일차결합을 만들어서 그 더해진 값이 활성화 함수의 값으로 들어가면 그 함숫값에 따라서 으로 나와서 반응을 하는지의 여부(output)를 볼 수가 있습니다.


    Schematically, it looks like the .\PICture below, and the output value is obtained through this process. In other words, when the product of an input ()  and a weight is found, it creates an output depending on the value of the weight function.


    지금 간단하게 인공신경망이 어떻게 움직이는지, 그에 필요한 용어들, Sigmoid function 그리고 활성화 함수(activation function), 그리고 받아들인 입력이 어떻게 활성화 함수를 통해서 출력으로 나오는 과정을 인공 뉴런에서 확인하였습니다.


    So far, we confirmed simply how the artificial neural network works, the terms necessary for it, the Sigmoid function and the activation function, and how the received input produces the output through the activation function.


    이어서 2차시에는 신경망의 학습을, 그리고 오차 역전파법(back propagation)에 대해서 자세하게 설명하도록 하겠습니다. 수고하셨습니다.


    Next, in the second lecture, we will study <how the neural network do study> and <the Backpropagation algorithm> in detail. Thank you.


    [13-2 pre]


    반갑습니다. 이제 13주차 2차시 강의에서는 <신경망의 학습>과 <오차역전파법>에 대해서 학습하겠습니다. 신경망은 입력층, 은닉층(Hidden layer), 출력층으로 구성되어 있는데, 입력층에서 신호를 받아 미리 부여된 가중치와 계산 후 주어진 활성화 함수를 거쳐 은닉층으로 전파되고, 같은 방식으로 출력층에 전달되어 예측 결과를 내보냅니다. 만일 주어진 데이터로부터 신경망을 이용하여 얻은 예측값과 미리 알고 있는 관측값 사이의 오차가 생기면, 이를 줄이는 방향으로 가중치를 점차 업데이트해 나가게 되는데, 이때 사용되는 대표적인 방법이 바로 오차 역전파법(back propagation)입니다.

    이번 시간에는 이에 대하여 자세히 학습하도록 하겠습니다.


    Welcome!!


    In the second session of this lecture, we will learn about <Learning of Neural Network> and <Backpropagation>.


    The neural network consists of input layer, hidden layers, and output layer. It receives a signal from the input layer and propagates it to the hidden layer through a given activation function after computation with a weight assigned in advance. The prediction result is transmitted to the output layer in the same way.


    If there is an error between the predicted value obtained using the neural network from the given data and the observed value known in advance, the weight is gradually updated in the direction of reducing it. Backpropagation and the gradient descent method are used.  In this session, we will learn more about this.


    [13주차 2강] 반갑습니다. 13주차 2강. <인공신경망>, 그 중에서 <오차 역전파법(Backpropagation)>에 대해서 학습하도록 하겠습니다.


    [Lecture 2, Week 13] Welcome to the week 13, 2nd class. We are going to learn about <Backpropagation> among <artificial neural network>.


    먼저 3절은 <신경망의 학습>으로, 우리가 그래프에서 보듯이 입력(input)이 있고 결과(output)가 나오는 사이인 그 중간에 어떤 과정을 거쳐서 결과가 생성되는지 살펴보겠습니다.


    First, this <Section 3> is about <Learning of Neural Networks>. As we can see in the graph, we will look what kind of process (a hidden layer)  were made between the input and the output result.


      신경망은 순환하지 않는 그래프로 연결된 뉴런의 모음을 모델화 한 것인데, 아래 그림에서 보듯이 왼쪽부터 입력층, 은닉층(Hidden layer), 출력층으로 구성되어 있습니다.


    A neural network is a model of a collection of neurons connected by a non-circulating graph. As shown in the figure below, it consists of an input layer, hidden layers, and an output layer from the left.


      입력층에서 신호를 받으면, 미리 부여된 가중치와 계산 후 주어진 활성화 함수를 거쳐 은닉층으로 전파되고, 같은 방식으로 그 다음 층으로 전해질 신호가 계산되고, 이런 과정을 반복한 후에 출력층에서 해당하는 결과를 내보냅니다.


    When a signal is received from the input layer, it is propagated to the hidden layer through a given activation function after calculation with the weight given in advance. In the same way, the signal is calculated and transmitted to the next layer, and after repeating this process, the output layer produces the corresponding result.


      이제 입력값과 정답을 알고 있는 데이터가 있다고 할 때, 우리가 수집한 데이터들, 이런 사람은 수치를 가지고 있다는 데이터를 가지고 있다고 할 때, 우리의 목적은 주어진 데이터를 잘 판단하도록 그에 맞는 신경망의 가중치를 찾는 것입니다. 이때 어떤 공식에 의해서 가중치를 계산하는 것이 아니라, 초기에 임의로 가중치를 부여해놓고, 주어진 데이터로부터 신경망을 이용하여 얻은 예측값과 미리 알고 있는 정답간의 오차를 줄이는 방향으로 가중치를 점차 업데이트해 나가는 것입니다. 이때 계층 간의 각각의 연결이 오차에 영향을 주는 정도에 비례하여 오차를 전달해주게 되는데, 대표적인 방법으로 오차 역전파법(Backpropagation)입니다.


    Now, assuming that there is data that knows the input value and the correct answer, our purpose is to find the weights of neural network to judge the given data correctly as much as we can. At this time, the weights are not computed by any formula, but the weights are given randomly at the beginning, and the weights are gradually updated in the direction of reducing the error between the predicted values using the neural network and the correct answers obtained from the given data. The representative method is the <Backpropagation(BP)>.


      이제 오차 역전파법, 즉 인공신경망을 학습시키기 위해서 사용하는 오차 역전파법에 대하여 알아보겠습니다.


     Let’s look at the Backpropagation method, which used to train artificial neural networks.


      이 설명은 제가 여러 문헌들을 참고하면서, 여러분들이 가장 이해하기 쉽도록 새로 쓴 설명입니다. 자세히 살펴보도록 하겠습니다.


    This explanation is a new and easy explanation that I have written for you to understand the Backpropagation algorithm.


      먼저, 데이터들이 있다면, 예를 들어, 병원에서 사람의 키와 나이, 성별 그리고 몸무게, 혈압, 체중, 하루에 식사하는 칼로리, 그리고 운동량 등 많은 데이터들이 있어서 대개 키가 몇이고 운동량이 얼마며, 식사를 얼마만큼 하는 사람이면 ‘체중은 얼마고, 혈압은 어느 정도고, 혈당은 어느 정도 되는 것이 정상이다’ 하는 식으로 데이터를 통하여 ‘이런 사람은 이렇다’라는 패턴은 알고 있습니다.


    First, if there are data, for example, there is a lot of data such as height, age, gender, weight, blood pressure, calories eaten per day, and exercise amount in hospital, the pattern of ‘This person is more like that person’ is known through the data such as ‘how much weight, how much blood pressure, and how much blood sugar is normal’.


      그런 데이터가 몇 십 년 동안 병원에 모여 현재 엄청나게 많은 몇 백만 명의 데이터들이 있습니다. 그 데이터들을 먼저 그냥 쓰기보다 아주 특수한 경우, 환자 중에서는 아주 눈이 띄게 특수한 경우들이 있으므로, 먼저 유의미 하지 않은 특수한 데이터들은 먼저 걸러냅니다. 이런 과정을 Data Cleaning이라고 합니다.


    That data has been gathered in hospitals for decades and now there is a huge amount of data from millions of people. Since there are very special cases, rather than just use the data, first, special data that are not meaningful are filtered out. This process is called a <data cleaning>.


      이제 데이터를 축, 축, 축, 축 하듯이 파일에 잘 맞게끔 organize 하여, 잘 정리하는 과정을 전처리 과정(Data Preprocessing)이라고 합니다. 전처리 과정을 잘 거쳐서, 필요한 모델을 만들 수 있는 기본 데이터들을 확보한 후에, 이 데이터(Set)를 학습 데이터(training data), 검증데이터(validation data), 테스트데이터로 일단 나눕니다.


    Now, the process of organizing data to fit into files is called a <Data Preprocessing>. After data preprocessing and securing the basic data to make the necessary model, divide this data set into training data, validation data, and a test data set.


      보통의 경우는 원래 데이터의 80%를 학습 데이터와 검증데이터로 사용하고, 20%를 테스트 데이터로 사용합니다. 그 후 학습데이터들을 이용해서 모델링을 하고, 검증데이터를 사용하여 모델링을 점점 정교하게 고쳐간 후에. 만들어진 모델을 가지고 테스트 데이터를 집어넣어, 잘 맞는지를 확인하는 과정을 거쳐 인공신경망에 맞는 함수, 즉 행렬을 찾아낸다고 볼 수 있습니다.


    In general, 80% of the original data is used as training and verification data, and 20% is used as a test data set. The modeling is performed using the training data, and it is refined more and more elaborately using the validation data. Then, it is checked whether it fits well by inserting test data. Through these process, we can find a function, that is, a matrix, that fits the given artificial neural network.

     

      이를 위하여 시작은 학습데이터로부터 입력과 관측값, 이것을 observed value라고 합니다.  이 둘을 먼저 입력 합니다. 입력과 observed value를 집어넣고, 대략 이걸 집어넣으면 이런 값이 나오게 인공신경망의 가중치를 임의로 설정합니다. 즉 적당한 함수로 임의의 행렬을 하나 제공해 주는 것입니다. 그리고 활성화 함수도, 우리가 ReLU 함수 또는 Sigmoid 함수를 정해주고 계수도 적당히 제공합니다. 예를 들어, 행렬의 경우 1, 1, ..., 1, 1 등으로 집어넣을 수 있습니다.

     

    To do this, first, enter input and observed value, then set the weight of artificial neural network arbitrarily. In other words, provides an arbitrary matrix for an appropriate function. Then, set the ReLU function or the Sigmoid function as an activation function and provide an appropriate coefficient. For example, for a matrix, you can put 1, 1, ..., 1, 1, etc.


      이제 이 모델에서 시작하여, 입력층에서 학습 데이터의 신호를 받으면, 미리 부여된 가중치를 받아서 계산 후 주어진 활성화 함수를 거쳐 은닉층으로 전파되고, 같은 방식으로 그 다음 층으로 전해질 신호가 계산됩니다. 이런 식으로 전파되어 Hidden layer를 다 거쳐서 출력층에서 해당하는 결과인 예측값(predicted value)을 얻게 됩니다. 이 예측값과 관측값 사이에는 오차가 생기게 될 것입니다. 이제 그 오차를 줄이는 과정을 거치게 되는데, 이때 우리가 앞에서 배운 경사하강법(gradient descent method)을 적용하여 이 오차를 최소화합니다. 이 답을 가지고 가중치와 행렬, 즉 함수를 수정합니다.


    Now starting with this model, when the input layer receives a signal from a training data, compute it with the weight and transmits to the hidden layer through a given activation function. In the same way, the signal pass through all hidden layers and we can get a predicted value at the output layer. There will be an error between this predicted value and the observed value. We minimize this error by applying the gradient descent method. With this answer, we modify the weights and matrices, i.e. functions.


      가중치를 수정하는 과정을 거쳐, 전체적인 오차가 최소화 될 때까지 ⓸~⓺ 과정을 반복합니다.


    After correcting the weights, repeat the ⓸~⓺ process until the overall error is minimized.


      그렇게 전체 과정을 거쳐서 우리가 optimal한 solution을 찾게 되면 신경망을 모델링한 셈이 됩니다. 이제 모델링을 하고 나서, 남겨둔 테스트 데이터를 이용하여 우리가 만든 신경망 모델이 잘 작동되는지를 테스트하게 되는 것입니다.


    If we find the optimal solution through the whole process, we have the neural network model. After modeling, check that the neural network model works well using the test data set.


      이 사이의 전체적인 과정 즉, Hidden layer 안에서 작동한 단계를 설명하는 이름이 바로 <오차 역전파법>입니다.


    The whole process operated within the hidden layers is called the <Backpropagation> algorithm.


      예를 들어, 100만개의 데이터를 다 사용하지 않고 그 중에서 카테고리로 나누어서 1/10만 데이터를 가지고도 모델링을 해 보고, 그 정확도를 더 높이려면, 데이터 숫자와 은닉층의 숫자를 조금씩 더 늘려가면서, 우리가 원하는 정도에 충분히 가까운 신경망을 모델링한다고 이해하시면 되겠습니다.


    For example, we see that we do find a model of a neural network with 1/100,000 of the original data at first. And if we want to increase the accuracy, then we do increase the number of training data and the number of hidden layers until we can have a model of enough accuracy as we need.


      이제 위의 <오차 역전파법 알고리즘>을 수학적으로 설명하도록 하겠습니다. 그리고  딥러닝(deep learning)에 대해서 설명해보겠습니다. 딥러닝이란 Hidden layer가 있는 인공신경망에서 Hidden layer가 2개 이상의 여러 개가 있는 경우에 layer가 깊어지니까 이런 인공신경망의 learning process를 딥러닝이라고 부릅니다.


    Let’s explain the <Backpropagation algorithm> and deep learning mathematically. Deep learning is a learning process of artificial neural network with more than two hidden layers.

     

      오차 역전파법에 의해서 각 계층에 전달된 오차를 바탕으로 가중치를 업데이트하는 방법은 이미 학습한 경사하강법을 오차함수를 최소화하는 문제에 적용한 것과 같습니다.


    The method of updating the weights based on error transmitted to each layer by the Backpropagation method is the same as applying the gradient descent method to the problem of minimizing the error function.


      예를 들어, 각 계층에서 입력신호 ()를 받아서 미리 부여된 가중치와 계산 후 주어진 활성화 함수로부터 출력된 값을 ( 아이 햇)이라고 합시다. 모자를 쓴 것 같죠. 그리고 실제 관측값(observed value, 정답에 해당하는 값)을 라고 한다면, 이 와 단순하게 출력된 사이에 오차가 존재합니다. 그럼 제곱 오차(squared error)를 다음과 같이 표시할 수 있습니다. 앞에서 보셨듯이,와 같이 와 사이의 차이를 노름(norm)을 이용해서 정의한 error function 이 정의됩니다.


    For example, let’s call the output value obtained from the activation function after computation with weights and input signal () received from each layer. And if the real observed value(the value corresponding to the correct answer) call , there is error between and . Then we can express a squared error as we can see here. The difference between and , an error function, is defined by using a vector norm.


      그러면 각 가중치를 업데이트하는 공식은 경사하강법으로부터 다음과 같이 주어집니다. 익숙한 함수입니다.


    Then the formula for updating each weight is given by the gradient descent method.


          에서 보듯이 하나씩 OLD - (에타 * 도함수) 를 계산하여 New 를 찾는 것입니다. 앞의 경사하강법을 설명할 때도 같은 식으로 를 로 정의하는 알고리즘을 사용했었습니다. 이때, 는 를 변수 에 관하여 편미분한다는 의미로 를 제외한 다른 변수는 모두 상수로 취급하여 미분하는 것입니다. 앞의 3장 ‘데이터와 최적화’에서 학습한 내용입니다.


    As we can see, , it is to find this NEW comes from the OLD - (eta*the derivative) one by one. In the previous description of the gradient descent method, the same algorithm was used to define as . At this time, means to differentiate with respect to the variable , that is, all other variables except for are treated as constants and differentiated. This is what you learned in Chapter 3, ‘Data and Optimization’.


      이제 아래와 같이 신경망의 각 계층에 전달된 오차로부터 가중치를 업데이트하는 방법에 관하여 살펴보면, 먼저 각 계층에 전달된 오차를 계산하고 그 다음 출력층에서의 오차를 계산하기 위하여, 은닉층에서 입력 신호 과 를 받아 가중치 들을 계산한 후에, 일차결합을 만듭니다. 그리고 Sigmoid 함수 를 거쳐 출력 과 를 얻었다고 가정합시다. 그러면 들이 들어가서 weight를 가지고 일차결합을 만들어, 출력층에서 결과가 나오게 되면 그것과 우리가 원래 가지고 있던 값과의 오차가 생기게 됩니다. 이 오차의 제곱들의 합을 error function으로 이해하고 그 오차들을 표현하면 다음과 같이 표현할 수 있습니다. weight를 이용하고, activation function을 사용합니다. 이제 은닉층에서의 오차를 계산해보는데, 은닉층에서는 출력층과 달리 목표하고자 하는 출력층에서의 관측값(측정값, 정답에 해당하는 값)이 따로 없으므로, 은닉층에서는 이 오차 에 해당하는 것이 없습니다. 이때 앞서 언급한 오차 역전파법을 활용하는 것입니다.


    Looking at the method of updating the weight from the error transmitted to each layer of the neural network, first, calculate the error transmitted to each layer. After calculate the weights with input signals and obtained from a hidden layer, make linear combination to compute an error of the next output layer. Let’s assume that we get output and through the Sigmoid function . When the result comes out from the output layer, there is an error between it and the observed value that we originally had. This error function, sum of squared errors, can be expressed as we can see here. Use the weights and the activation function. When calculating this error in this hidden layer, since there is no observed value, unlike the output layer, there is no corresponding error . In this case, the aforementioned Backpropagation method is used.


      신경망에서 오차가 생겼다는 것은, 입력신호가 입력층으로부터 은닉층을 거쳐 최종 출력층으로 전파될 때, 은닉층에서의 오차가 반영된 결과라고 볼 수 있으므로 계층(layer) 사이에 각각의 연결이 오차에 영향을 주는 정도, 즉 가중치에 비례해서 오차를 역으로 전달해주면 되는 것입니다. 그러니까 같은 식으로 은닉층에서 이 그림같이 출력이 나오면 우리가 예상하는 값과의 오차를 계산해서 다음 단계로 돌려보내주는 것입니다.


    The error in a neural network when the input signal propagates from the input layer through the hidden layer to the output layer can be seen that the error in the hidden layer is reflected. Thus, we can transmit the error inversely in proportion to the degree of influence on error, that is, the weight. So, in the same way, if the output is shown in this .\PICture from the hidden layer, we compute the error between the expected value and apply it to the next step.


      그래서 출력층의 첫 번째 노드에서 오차 을 얻었다고 가정하고, 은닉층의 첫 번째 노드와 가중치 , 두 번째 노드와 가중치 으로 연결되니까 가중치가 크면 오차에 미치는 영향도 클 것이므로, 가중치에 비례하여 오차를 은닉층에 전달해줍니다. 예를 들어, 를 은닉층의 첫 번째 노드에 전파된 오차라 하면 다음과 같이 표현할 수 있습니다.  는 다음과 같이 표현할 수 있습니다. 마찬가지로 도 구할 수 있습니다. (화면의 수식을 보세요)


    So, assuming that the error was obtained at the first node of output layer, it is connected to the first node of the hidden layer with the weight and the second node of the hidden layer with the weight . If the weight is large, the effect on the error will be large, so the error is transferred to the hidden layer in proportioin to the weight. For example, if is an error propagated to the first node of the hidden layer, it can be expressed as this. Also, in the same way, can be expressed as this.


             ,    


                                                             <---화면에 있으니 자막에는 없어도 됨)


    위 두 식에서 분모를 제외하면 다음과 같이 행렬을 이용하여 표현할 수 있습니다. 이때, 분모를 없애게 되면 원 식과 비교했을 때, 일정 부분의 비율이 사라지는 정도의 효과가 있게 되는데, 다음 웹 사이트를 보시면 아시겠지만, 분모가 있는 경우와 실제로는, 우리가 얻은 결과가 거의 차이가 없이 잘 작동되므로, 이 상수배는 우리가 찾는 optimal solution에 큰 영향을 미치지 않기 때문에 이렇게 단순화 시켜서 선형연립방정식 문제로 바꾸어서 앞에서 배운 optimal solution을 구하는 알고리즘을 활용해서 인공신경망, 특히 Hidden layer, 그리고 Backpropagation 오차 역전파법을 활용할 수 있는 것입니다.  (화면의 수식을 보세요)


    If the denominator is removed from the above two equations, it can be expressed using a matrix as follows. At this time, if the denominator is removed, the ratio of a certain portion disappears when compared to the original equation. As we can see from the following website http://matrix.skku.ac.kr/math4ai-intro/W13/, there is almost no difference between the case where the denominator exists and the result we have obtained. Since it does not have a big effect on optimal solution, we may simplify it and turn it into a linear system problem. Then, we use the artificial neural network, especially the hidden layer, and the Backpropagation method.


                         

                                     <---화면에 있으니 자막에는 없어도 됨)


      위는 웹 자료인 ‘What is Backpropagation really doing?’ 설명에서 볼 수 있듯이 그 내용을 쉽게 설명해드린 것입니다.


    The above is an easy updated explanation of the content as we can see in the explanation of ‘What is Backpropagation really doing?’


    두 번째, 출력층에서 얻은 오차로부터 은닉층과 출력층 사이의 가중치를 업데이트합니다. Sigmoid 함수의 성질과 연쇄법칙으로부터 을 구하면 다음과 같이 구해질 수 있습니다. 앞에서 보았듯이 위 함수의 도함수를 구하면 Sigmoid function의 성질에 따라 도함수가 이렇게 쉽게 구해집니다. 직접 계산하지 않으시더라도 Sigmoid function의 성질 때문에 이렇게 쉽게 나온 결과를 얻을 수 있다는 것은 바로 이해하시기만 하면 됩니다.


    Second, we update the weights between the hidden layers and the output layer from the error obtained at the output layer. From the properties of the Sigmoid function and the chain law, we can find as follows. As we have seen before, if we find the derivative of the above function, the derivative can be easily obtained according to the properties of the Sigmoid function. We really do not have to compute it by ourselves, we can get the result easily because we already know the nature of the Sigmoid function.


      같은 방법으로 다른 가중치에 대해서도 쉽게 얻을 수가 있고, 여기서 가중치 는 은닉층의 번째 노드와 출력층의 번째 노드만 연결되어 있으므로, 를 에 대하여 편미분하면 다음과 같이 간단하게 표현되는 걸 확인할 수 있습니다. 여기서 는 은닉층의 번째 노드에서의 입력이고, 는 출력층의 번째 노드에서의 출력이며, 로 이것은 출력층의 번째 노드에서의 오차가 됩니다.


    In the same way, it is easy to obtain other weights, and here, since the weight is connected only with the th node of the hidden layer and the th node in the output layer, can be expressed simply as follows. Here, is an input from the th node in the hidden layer, is an output from the th node in the output layer and is an error at the th node in the output layer.


      이제 경사하강법을 이 오차에 적용하여 은닉층과 출력층 사이의 모든 가중치를 업데이트합니다. 이게 경사하강법의 알고리즘으로 같은 내용입니다.


    Then, we just apply the gradient descent method to this error to update all weights between the hidden layers and the output layer. This is the same contents as the gradient descent algorithm did.


      그것을 반영한 후에, 은닉층에 전달된 오차로부터 입력층과 은닉층 사이의 가중치를 업데이트합니다. 기본적으로는 은닉층과 출력층 사이의 가중치를 업데이트하는 방법으로 계산하면 되는데, 논의의 편의상 기호를 그대로 사용하면, 은닉층 안에서도 마찬가지로 같은 알고리즘인 오차 역전파법(Backpropagation) 알고리즘이 돌아가는 걸 이해하실 수 있습니다.


    After reflecting it, update the weights between the input layer and the hidden layer from the error passed to the hidden layer. Basically, we can compute it by updating the weights between the hidden layer and the output layer. For convenience, if we use the symbol as it is, we can see that they are the same algorithm, the Backpropagation, works in the hidden layer.


      여기서 는 입력층의 번째 노드에서의 입력, 는 은닉층의 번째 노드에서의 출력, 는 은닉층의 번째 노드에 전파된 오차이므로 다음과 같이 는 이렇게 표현할 수 있고, 경사하강법을 적용하여 또 모든 가중치를 다음과 같이 업데이트할 수 있습니다.


    Here, is an input from the th node in the input layer, is an output from the th node in the output layer, is an error passed to th node. The can be expressed like here in, and all weights can be updated as follow by applying the gradient descent method.


      신경망은 모델이 제대로 예측할 때까지 많은 데이터를 필요로 하고, 또한 입력층과 출력층 사이에 많은 은닉층을 둘 수도 있습니다. 이와 같이 은닉층이 여러 개 있는 인공신경망을 심층신경망(deep neural network)이라고 부르며, 심층 신경망을 학습하기 위한 기계학습(Machine learning)을 딥러닝(deep learning)이라고 부르는 것입니다. 은닉층이 여러 개 있는 경우에도 마찬가지로 오차 역전파법을 이용하여 그 이전 층에서 전파된 오차로부터 경사하강법을 적용하여 가중치를 업데이트할 수 있습니다. 이에 관련된 다양한 자료들이 있으므로 참고하시고, 과제로 인공신경망과 오차 역전파법에 대해서 개인 수준에 따라 이해하신 대로 요약하시면 됩니다.


    Neural networks require a lot of Data until the model can predict properly, and may also have many hidden layers between the input and the output layers. Such an artificial neural networks with several hidden layers are called a <Deep Neural Network>, and a machine learning for learning deep neural network is called the Deep Learning. If there are a multiple hidden layers, we can also use this Backpropagation method to update the weights by applying a gradient descent method from the error propagated in the previous layer. There are various data related to this, so we can refer to it. For homework, summarize the artificial neural network and the Backpropagation method as you understand.


    이제 오늘 배운 내용을 복습합시다.

    Now let's review what we learned today.


    [W 13 Review] [W 13 Review]


    13주차 강의에서 다룬 인공신경망(Artificial Neural Network)은 신경계의 기본 단위인 사람의 뉴런을 모델화 한 것으로, 이미지 인식, 글자 인식 등 다양한 업무를 사람이 하듯이 잘 수행하고 있습니다. 이번 13주차에서는 신경망의 작동 원리, 특히 <오차 역전파법(Backpropagation)>의 알고리즘을 자세히 설명하였습니다. 그리고 신경망과 신경망의 작동 원리를 학습하였고, 그 안에 행렬이 사용되고 activation function으로 Sigmoid 함수가 사용된 것과 오차 역전파법의 알고리즘, 여기에 경사하강법(Gradient Descent Method)이 어떻게 사용되었는지 설명하였습니다.


    The artificial neural network was a mathematical model of a human neuron, the basic unit of nervous system. It performs various tasks such as image recognition and text recognition as well as human face recognition. In this week, we studied how the neural network works, especially the algorithm of <Backpropagation> in detail. In addition, we practiced a neural network and how it works, a matrix was used in it, this Sigmoid function was used as an activation function and we learned how this Backpropagation algorithm and the gradient descent method was used in it.


      이상으로 이번 13차 강의를 마치고 다음 시간에 MNIST 데이터를 이용한 손 글씨 인식에 대해서 배우고 실습해보도록 하겠습니다.  수고 많이 하셨습니다.


    That’s the end of today's lecture. Next time, we will learn and practice handwriting recognition using MNIST Data set.


     


    Week 14. Introductory Mathematics for Artificial Intelligence

                인공지능을 위한 기초수학 입문


    14. MNIST 데이터 숫자인식 실습 157

    14.1 인공신경망을 활용한 손 글씨 숫자 인식 사례 

     - 과제 (열린문제)-  166

      - Final PBL 보고서  167


    [14-1 pre]

    여러분, 정말 반갑습니다.

    Everyone, nice to meet you.


     이제 K-MOOC [인공지능을 위한 기초수학 입문] 14주차 마지막 주 첫 수업입니다.

    This is the first lecture in the final week of K-MOOC [Introduction to Basic Mathematics for Artificial Intelligence class.


    지금까지 같이 해주셔서 정말 감사합니다.

    Thank you so much for being with me so far.


    이번 시간에는 지금까지 배운 모든 지식을 직접 활용해보도록 합니다.

    In this lecture, we will try to use all the knowledge we have learned so far.


    14주 차에는 MNIST 데이터 set으로부터 우리가 손으로 쓴 숫자를 인공지능이 어떻게 인식해서 그 숫자가 무엇에 가장 가깝다는 것을 알려주는 과정을 실제 확인해보도록 하겠습니다. 이제 <인공신경망을 활용한 손글씨 숫자 인식 사례>에 대해서 실습해보겠습니다.

    We'll see how the AI recognizes handwritten numbers from the MNIST data set and tells us what the numbers are closest to. Now, let's practice on <Handwriting Number Recognition Using Artificial Neural Network>.


    과거 우체국에서는 손으로 적은 우편번호를 인식하고 분류하며 배달하는 것에 많은 인력을 투입했었습니다.

    In the past, post offices spent a lot of work on recognizing handwritten zip codes, sorting them, and then deliver.


    이 절에서는 지난 시간에 학습한 인공신경망을 활용하여 MNIST 데이터 셋에 있는 숫자 이미지들을 활용하여, 우리가 손으로 쓴 글자를 인식한 후, MNIST 데이터 셋의 숫자 중 어느 숫자에 가장 근사한 숫자를 찾아서, 우리가 쓴 우편번호가 무엇에 가장  가까운지를 파악하는 과정을 실습합니다.


    In this section, we see how the artificial neural network recognizes the handwritten numbers in the MNIST data set and finds the number what is the closest to. We practice the process of figuring out what the zip code we wrote is closest to.


    이 정보를 이용하여 편지봉투들을 자동으로 분류해서, 그 지역을 담당한 우편배달부가 들고 바로 업무를 수행할 수 있도록 해주는 과정을, 사람이 아니라 인공지능이 대신 해주는 그 과정과 원리에 대해서 학습하겠습니다.

    Using this information, we will learn about the process and principle how artificial intelligence instead of humans does the process of automatically classifying envelopes, so that the postal delivery department in charge of the area can carry out their work right away.


    1차시에는 <인공지능을 활용한 손글씨 숫자 인식 사례>를 설명하고 그에 관련된 코드를 소개합니다. 실제 손으로 쓴 숫자를 인식하고 데이터들과 비교하여, 데이터 셋에 있는 숫자들과 우리가 손으로 쓴 글자를 비교하여, 8 또는 3으로 인식되는 과정과, 우리가 쓴 숫자 7이 인공신경망을 거쳐서 7에 가장 가까운 숫자를 썼다는 output을 우리에게 알려주게 하는 그 전체 과정과 그 과정에 쓰이는 오차 역전파법 그리고 오차 역전파법을 중간과정에서 계속 update 하는 동안 경사하강법이 활용되는 것을 실제 실습을 통하여 학습하도록 하겠습니다.


    In the first lecture, we explain <Examples of Handwritten Number Recognition Using Artificial Intelligence> and introduce related codes. We can understand the process of AI’s recognizing actual handwritten numbers and comparing them with data, comparing the numbers in the data set with the handwritten letters. Back propagation and the gradient descent method are the main algorithms to update the weights in Artificial Neural Network.


    [14주차 1강] 여러분 반갑습니다. 고교생과 일반인들을 위한 K-MOOC [Introductory Mathematics for Artificial Intelligence] 과목의 마지막 주 14주차 강의입니다.


    [Lesson 1, Week 14] We are in the 14th week. This will be the last week of the K-MOOC [Introductory Mathematics for AI] course. Congratulations!!


    이제, 앞에서 배운 그 모든 지식이 활용되는 예를 소개하겠습니다. 여기서는 MNIST라고 불리는 숫자인식 이론과 실제를 학습하도록 하겠습니다.


    Now, let's look at an example of all the knowledge we've learned earlier being utilized. Here we learn the theory and practice of a handwritten number recognition with the MNIST dataset and ANN.


    MNIST 데이터셋을 이용한 실습실은 14주차에 만들어 놨으니까, 보시면서 직접 실습을 하도록 하겠습니다. http://matrix.skku.ac.kr/math4ai-intro/W14/


    The practice Lab http://matrix.skku.ac.kr/math4ai-intro/W14/ using the MNIST dataset was made for you for this Week 14.


    1절은 <인공신경망을 활용한 손 글씨 숫자 인식>에 대한 내용입니다. 같은 원리가 여러 가지 QR 코드를 인식하는 것부터 시작해서 bar코드를 인식하는 것 등에 적용되는 것입니다.


    The <Section 1> is about <Hand Writing Numbers detection using Artificial Neural Network>. The same principle can be applied to recognize various QR codes and bar codes etc.


    과거 우체국에서는 손으로 적은 우편번호를 인식하고 분류하고 배달하는 것에 많은 사람들이 투입되기도 했었습니다. 이 절에서는 지난 시간에 학습한 인공신경망을 활용하여 컴퓨터가 손으로 적은 숫자 이미지를 인식하는 원리에 대하여 학습합니다.


    In the past, many people were involved in [Recognition of Handwritten (postal) ZIP Codes in a Postal Sorting system]. In this section, we learn about how Computer/AI do recognize Handwritten ZIP Codes by artificial neural networks.


     지금은 우체국에서 바로 인공지능이 우편번호를 인식해서 저절로 분류를 해 줌으로서 많은 인력을 대체하고 있습니다. 여기 이미지들을 보시면 0, 1, 2, 3, 4, 5, 6, 7, 8, 9라는 글자를 사람에 따라서 다양하게 쓰고 있습니다.


    Now, artificial intelligence (Robot) is replacing a lot of manpower at the post office by recognizing ZIP code and sorting them out by themselves. If we look at the images here, the letters 0, 1, 2, 3, 4, 5, 6, 7, 8, 9 can be written in various ways by each person.


    위의 그림과 같이 손으로 적은 숫자 이미지들이 인공신경망을 활용하려면 우선적으로 필요합니다. MNIST는 Modified National Institute of Standards and Technology의 약자인데, 이곳이 제공하는 데이터베이스에는 가로, 세로가 각각 28 픽셀(pixel)인 흑백의 숫자 이미지가 약 70,000개가 들어있습니다. 그 중 60,000개는 학습용 데이터이고, 10,000개는 테스트용 데이터입니다. 각 데이터에는 손으로 쓴 0부터 9까지의 숫자 이미지와 그 레이블(label), 즉 실제 숫자에 대한 정보가 담겨 있습니다. 기계 학습의 성과를 벤치마킹 하는데 가장 널리 사용되는 MNIST 데이터는 아래 웹 사이트에서 확인할 수 있습니다. 이제 MNIST 데이터의 숫자 이미지를 실제로 확인해봅시다.


    In order to utilize the artificial neural network as shown in the .\PICture above, a hand-written numeral images (number) are needed first. MNIST is an abbreviation for <Modified National Institute of Standards and Technology>, whose database contains approximately 70,000 numeric images in black and white, 28 pixels wide by 28 pixels tall. Of these, we devide it as 60,000  learning data and 10,000 test data. Each data contains a hand written image (with a label) of numbers from 0 to 9. The most widely used MNIST data for benchmarking the performance of machine learning can be found on the Web site below. Now let's actually check out the number images of the MNIST data.


    여기서는 https://github.com/freebz/Make-Your-Own-Neural-Network 가 말하듯이 Neural Network을 실제로 만들어보자는 제목으로 깃허브(github) 페이지에 공개된 Python 코드를 우리에게 맞게끔 일부 수정하여 자세한 설명을 줄 마다 달아서 여러분이 이해하기 쉽게 최초로 공개한 것입니다.


    Here, we have modified the Python code to Sage, which was released on a GitHub page in https://github.com/freebz/Make-Your-Own-Neural-Network.

    As the title 'Make-Your-Own-Neural-Network' show, it works with our code that fits us. We added detailed explanations to make it easier to be understood.


    이 알고리즘을 천천히 이해하고, 이어서 직접 실습을 해보도록 합시다. 먼저 첫 부분에는 필요한 라이브러리로, numpy 와 matplotlib 등을 부른 후 데이터 파일을 불러옵니다. (화면의 코드를 보세요)


    Let's read the code and try to understand the algorithm. Then we can practice it in the Lab http://matrix.skku.ac.kr/math4ai-intro/W14/.  


    The first part of the code calls <numpy and matplotlib>, and some needed library.


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    # 필요한 라이브러리 불러오기 # Load required library

    import numpy

    import matplotlib.pyplot

    import csv

    import urllib.request

    import codecs


    # mnist 학습 데이터인 웹에 있는 csv 파일을 리스트로 불러오기

    # Load csv file from web, mnist training data, into list

    url = 'https://media.githubusercontent.com/media/freebz/Make-Your-Own-Neural-Network/master/mnist_dataset/mnist_test_10.csv'


    ---------------------------------------------------------------------

                                                             <---화면에 있으니 자막에는 없어도 됨)


    그 다음 이 import 명령어를 이용해서 mnist 학습 데이터인 웹에 있는 csv 파일을 불러옵니다. 이 url 을 주고, media 주소에 mnist_test_10.csv 주어 그 파일을 불러옵니다. 그 후 이 데이터를 열고 csv 파일을 읽은 후에 리스트로 만들고 training data 하나로 정의합니다. 그 다음 28 x 28 행렬로 바꿉니다. 이것이 이미지의 픽셀 크기였습니다. 그리고 이 행렬을 시각화하고, 이 명령어를 통해서, 실행하면 우리가 얻은 이미지가 이제 여기 foo.png 파일로 저장된 것입니다. (화면의 코드를 보세요)


    Next, we used the import command to load the csv file from the given URL on the web, which is the MNIST learning data. Give this URL and give the media address mnist_test_10.csv to load the file. Then open this data, read the csv file, make a list, and define it as a training data. Then the code transforms this training data to a 28x28 matrix. This is the pixel size of the image. And when we visualize this matrix through this command, the image we get is saved as a foo.png file.


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    # 웹 데이터 열기 # Open web data

    response = urllib.request.urlopen(url)


    # csv 파일 읽기 # read csv file

    training_data_file = csv.reader(codecs.iterdecode(response, 'utf-8'))

    training_data_list = list(training_data_file)  # 리스트로 만들기

    all_values = training_data_list[0]  # 데이터 한 개


    # 저장된 숫자가 무엇인지 파악하기 # Figure out what the stored number is

    print(all_values[0])


    # 28 x 28 행렬로 바꾸기 # Convert to 28 x 28 matrix

    image_array = numpy.asfarray(all_values[1:]).reshape((28,28))


    # 행렬을 이미지로 시각화하기 # Visualize the matrix as an image

    matplotlib.pyplot.imshow(image_array, cmap = 'Greys', interpolation = 'None')


    # 이미지를 foo.png로 저장 # Save image as foo.png

    matplotlib.pyplot.savefig('foo.png')


    # 웹 데이터 닫기 # Close web data

    response.close()

    ---------------------------------------------------------------------

                                                             <---화면에 있으니 자막에는 없어도 됨)


    이렇게 명령어를 복사해서, 이 실습실에서 구현을 시키면, 다음과 같이 7이라는 숫자가 나오고 png 파일로 이미지가 생성됩니다.


    When we copy and execute the command in our lab, the number 7 appears and the image is created as a png file as follows.


    이때 training_data_list[0]에 저장된 숫자가 7이었는데, 이 이미지를 클릭하면 다음과 같이 숫자 7을 이미지로 보여주는 것입니다. 같은 방법으로 training_data_list[1] 로 바꿔주면 저장된 숫자가 2이므로, 거기에는 2라는 이미지가 보여집니다. 즉 다른 이미지로 바꿔주면 답이 다른 숫자로 나오게 되는 것입니다.


    At this time, the number stored in training_data_list[0] was 7, and when we click on this, the number 7 is shown as an image as we see now.


    If we change this number [0] to training_data_list[1] in the same way, the stored number is 2, so the image 2 will be seen. In other words, if we choose another (number) image, ANN will tell us what was that number.


                      all_values = training_data_list[1]  # 데이터 한 개


    그럼 이제 숫자에 있는 이미지가 어떻게 행렬로 변환되는지를 한번 봅시다. 이 이미지를 가로 28픽셀, 세로 28픽셀의 크기로 나누고, 28 곱하기 28인 784개의 픽셀로 밝기를 조사하고, 밝기는 0을 검은색으로, 255를 흰색으로 설정하는데, 이 숫자, 픽셀에는 이미지의 밝은 정도를 나타내는 숫자가 들어 있다고 볼 수 있습니다. 각 픽셀의 위치와 밝기를 예를 들어, 가로 12, 세로 6번째 픽셀의 밝기가 82이라는 의미는 (12, 6, 82)로 표현하면 되겠습니다. 그래서 아래와 같이 숫자 이미지를 행렬로 나타낼 수 있는 것입니다. 여기에서 보면 검은색은 다 0으로 표시되고 색깔 있는 곳에 82가 쭉 찍혀서 8이라는 글자가 보여집니다. 이 행렬 데이터가 이 글자를 의미하는 것으로 이해할 수 있습니다.


    Let's take a look at how the images of a number was converted into a matrix. Divide the image into sizes of 28 pixels wide by 28 pixels tall. Now we have a image/matrix of 784=28x28 pixels. And we classify the brightness of each pixel with the natural numbers. The brightness of 0 is set to black and the number 255 to white, and then we have a 28x28 matrix whose entries indicating its brightness. For example, “the horizontal 12th and vertical 6th pixel and brightness is 82” would be expressed as (12, 6, 82). Thus, a number image can be represented in a matrix as shown below.


    Here, all black image is a 28x28 zero matrix. A colored area is marked with the number 82.  Now the letter 8 is shown below. It can be understood that this matrix data means a hand written letter.

     

    이제 신경망을 통하여 숫자 이미지를 인식하는 과정을 살펴보면, 학습용 데이터의 숫자 이미지 한 개는 28 × 28의 행렬이므로, 이것을 다음과 같이 784 × 1의 벡터로 변환을 시킵니다.


    Now, if we look at the process of recognizing the image of a number through an ANN, the image of a number in our learning data is a matrix of 28×28, which converts it into a 784×1 vector.


    이 벡터의 각 성분은 입력신호가 되어 입력층의 노드 로 입력이 되고 신경망을 거쳐 출력층의 신호를 내보내는데, 이때 출력층의 신호가 가장 큰 노드(node)의 레이블(label)로 숫자를 판단하게 되는데, 예를 들어, 출력층의 신호가 가장 큰 노드의 레이블이 7이면 숫자 7에 가까운 글씨였다고 인식을 하게 되는 것 입니다. 숫자 7이라는 28 × 28 픽셀이 벡터로 784 × 1의 벡터로 전환되어 입력층에 전달이 되면 신경망 안에서 Hidden layer를 거치면서 오차 역전파법 알고리즘을 거쳐서 0부터 9까지의 숫자 중에서 7에 가장 가까운 숫자 그런 글씨였다고 인식을 하는 것입니다.


    Each component of this vector is an input signal and will be sent to node of the input layer. Then the output layer through the ANN will give us an answer.


    At this point, the signal at the output layer is determined by the label of the largest node. In the earlier example, the output layer with the largest node was 7, then it is perceived as the given number was more close to the number 7. When a 28×28 pixels image of the number 7, were converted to a 784 × 1 vector and delivered to the input layer, the number closest to 7 will be recognized in this neural network through the Hidden layers and the Backpropagation algorithm.


    이를 실제 정답과 비교하여 오차가 생기면, 경사하강법과 오차 역전파법을 활용하여 신경망의 가중치를 조정하는 것입니다. 이렇게 모든 학습 데이터를 여러 번 사용하여 학습이 끝나면 테스트 데이터를 이용하여 신경망의 성능을 확인하게 됩니다.


    If there is an error by comparing this with the actual correct answer, the weight of the neural network is adjusted using the gradient descent method and the Backpropagation algorithm. After all these learning data was used (when the learning process is completed), the remaining test data will be used to determine the performance of our neural network.


    다음은 입력, 은닉, 출력 계층의 신경망으로, MNIST 데이터 세트 일부를 학습시켜 테스트용 숫자 이미지를 잘 인식하는지 보여주는 코드입니다. 여기서 학습 데이터는 100개, 테스트 데이터 10개를 아래 공개된 자료로부터 사용하였습니다.


    The following is a neural network of the input, hidden layers, and output layers. This code shows how ANN worked when we use 100 learning data and 10 test data in it.


    숫자를 인식하기 위한 신경망의 입력층은 784개의 노드와 출력층은 0부터 9까지의 숫자를 나타내는 10개의 노드로 구성되며 은닉층의 노드의 개수에 관하여 정해진 규정은 따로 없으나 이 절의 예시코드에서는 100개로 하였습니다. 그 코드를 한번 자세히 살펴보겠습니다.


    The input layer of the neural network for recognizing numbers consists of 784 nodes and the output layer consists of 10 nodes representing numbers from 0 to 9. There are no specific regulations concerning the number of nodes in the hidden layer, but the number of nodes in this section was set at 100. Let's take a closer look at the lines in the code.



    계층이 3개인 신경망으로 MNIST 데이터를 학습하는 코드를 설명한 것입니다. numpy를 import 하고, 시그모이드 함수를 사용하기 위해서 scipy.special를 불러오고, 행렬을 시각화하기 위한 라이브러리를 import하고, 신경망의 클래스를 neuralNetwork 으로 정의하고, 신경망을 초기화합니다. input, hidden, output, learningrate, 그 다음에 가중치 행렬 wih와 who를 정의하고, 배열 내 가중치는 w_i_j로 표기하며 w_i_j는 노드 i에서 다음 계층의 노드 j로 연결되는 것입니다. 학습률(learning rate)을 정의해주고, activation function으로는 시그모이드 함수를 정의해줍니다.  (화면의 코드를 보세요)


    This code explains that it learns from the data with a three-layer neural network. It starts to import a 'numpy' and load 'scipy.special' to use the 'sigmoid function'. It imports library for visualizing matrices, it defines 'Class of neural network' as 'neuralNetwork' and it initializes a 'Neural network'. Then it defines 'input, hidden layer, output, learning rate, and weight matrix, and the weights in the array are denoted by w_i_j which means connecting from node i to node j, in the next layer. It defines the learning rate and the activation function was defined as a Sigmoid function.


     ---------- http://matrix.skku.ac.kr/KOFAC/ --------------------------

    # 3계층의 신경망으로 MNIST 데이터를 학습하는 코드

    # Code to learn MNIST data with 3 layers of neural network

    import numpy

    # 시그모이드 함수 expit() 사용을 위해 scipy.special 불러오기

    # Call scipy.special to use the sigmoid function expit()

    import scipy.special

    # 행렬을 시각화하기 위한 라이브러리

    # Library for visualizing matrices

    import matplotlib.pyplot


    #신경망 클래스의 정의

    #Definition of neural network class

    class neuralNetwork:

        

        # 신경망 초기화하기

    # Initialize the neural network

        def __init__(self, inputnodes, hiddennodes, outputnodes, learningrate):

            # 입력, 은닉, 출력 계층의 노드 개수 설정

            self.inodes = inputnodes

            self.hnodes = hiddennodes

            self.onodes = outputnodes

            

           # 가중치 행렬 wih와 who

            # 배열 내 가중치는 w_i_j로 표기. 노드 i에서 다음 계층의 노드 j로 연결됨을 의미

           # w11 w21

           # w12 w22 등

    # Weight matrix wih and who
    # Weight in the array is expressed as w_i_j.It means connecting from node i to node j in the next layer.
    # w11 w21
    # w12 w22 etc

            self.wih = numpy.random.normal(0.0, pow(self.hnodes, -0.5), (self.hnodes, self.inodes))

            self.who = numpy.random.normal(0.0, pow(self.onodes, -0.5), (self.onodes, self.hnodes))

            

            # 학습률

    # Learning rate

            self.lr = learningrate

            # 활성화 함수로는 시그모이드 함수를 이용

    # Use sigmoid function as activation function

            self.activation_function = lambda x: scipy.special.expit(x)

            

            pass


    ---------------------------------------------------------------------

                                                             <---화면에 있으니 자막에는 없어도 됨)


     그 다음에 신경망을 학습시키기 위해서 입력 리스트를 2차원의 행렬로 변환시키고 은닉 계층으로 들어오는 신호를 계산하여 은닉 계층에서 나가는 신호를 구하고, 최종 출력 계층으로 들어오는 신호를 계산하여 최종 출력 계층에서 나가는 신호를 찾고, 그 error 출력 계층에서의 오차 (실제 값 - 계산 값)를 계산한 후에 은닉 계층의 오차는 가중치에 의해 나뉜 출력 계층의 오차들을 재조합해 계산합니다. 그래서 은닉 계층과 출력 계층 간의 가중치 업데이트하고, 이 코드를 이용해서 입력 계층과 은닉 계층 간의 가중치 업데이트합니다. (화면의 코드를 보세요)


    For the learning process, the neural network converts a given data to a two-dimensional matrix and then computes the input signal and output signal through the hidden layers.


    In this process <the Backpropagation algorithm> is used over and over again until it gives us the final output.  (See the code below)


    ---------------------------------------------------------------------

     # 신경망 학습시키기

    # Train neural network

        def train(self, inputs_list, targets_list):

            # 입력 리스트를 2차원의 행렬로 변환

    # Convert input list to 2-dimensional matrix

            inputs = numpy.array(inputs_list, ndmin=2).T

            targets = numpy.array(targets_list, ndmin=2).T

            

            # 은닉 계층으로 들어오는 신호를 계산

    # Compute the signal coming into the hidden layer

            hidden_inputs = numpy.dot(self.wih, inputs)

            # 은닉 계층에서 나가는 신호를 계산

    # Compute the signal going out of the hidden layer

            hidden_outputs = self.activation_function(hidden_inputs)

            

            # 최종 출력 계층으로 들어오는 신호를 계산

    # Compute the signal coming into the final output layer

            final_inputs = numpy.dot(self.who, hidden_outputs)

            # 최종 출력 계층에서 나가는 신호를 계산

    # Compute the outgoing signal from the final output layer

            final_outputs = self.activation_function(final_inputs)

            # 출력 계층의 오차는 (실제 값 - 계산 값)

    # The error of the output layer is (actual value-calculated value)

            output_errors = targets - final_outputs

            # 은닉 계층의 오차는 가중치에 의해 나뉜 출력 계층의 오차들을 재조합해 계산

    # The error of the hidden layer is calculated by recombining the errors of the output layer divided by the weight.

            hidden_errors = numpy.dot(self.who.T, output_errors)

            # 은닉 계층과 출력 계층 간의 가중치 업데이트

    # Update weights between hidden layer and output layer

            self.who += self.lr * numpy.dot((output_errors * final_outputs * (1.0 - final_outputs)), numpy.transpose(hidden_outputs))

            # 입력 계층과 은닉 계층 간의 가중치 업데이트

    # Update weights between input layer and hidden layer

            self.wih += self.lr * numpy.dot((hidden_errors * hidden_outputs * (1.0 - hidden_outputs)), numpy.transpose(inputs))

            

            pass


    ---------------------------------------------------------------------

                                                             <---화면에 있으니 자막에는 없어도 됨)


    그렇게 하고 신경망한테 물어봅니다. 다음과 같이 query를 하는데, 입력 리스트를 2차원 행렬로 변환하여서, 은닉 계층으로 들어오는 신호를 계산하고, 은닉 계층에서 나가는 신호를 계산해서 최종 출력 계층으로 들어오는 신호를 계산으로부터 최종 출력 계층에서 나가는 신호를 찾으라는 프로그램을 다 갖다 놓고, 이제 실제 데이터를 주고 실행을 시켜봅니다. (화면의 코드를 보세요)


    Now practice the given code to see how it works.  (See the code below)


    ---------------------------------------------------------------------

    # 신경망에 질의하기

    # Querying the neural network

        def query(self, inputs_list):

            # 입력 리스트를 2차원 행렬로 변환

    # Convert input list to 2D matrix

            inputs = numpy.array(inputs_list, ndmin=2).T

            # 은닉 계층으로 들어오는 신호를 계산

    # Compute the signal coming into the hidden layer

            hidden_inputs = numpy.dot(self.wih, inputs)

            # 은닉 계층에서 나가는 신호를 계산

    # Compute the signal going out of the hidden layer

            hidden_outputs = self.activation_function(hidden_inputs)

            # 최종 출력 계층으로 들어오는 신호를 계산

    # Compute the signal coming into the final output layer

            final_inputs = numpy.dot(self.who, hidden_outputs)

            # 최종 출력 계층에서 나가는 신호를 계산

    # Compute the outgoing signal from the final output layer

            final_outputs = self.activation_function(final_inputs)

            

            return final_outputs

    ---------------------------------------------------------------------

                                                             <---화면에 있으니 자막에는 없어도 됨)


    첫째, input_nodes를 784, hidden_nodes는 이번에는 100으로 하고, 출력 nodes 수는 10 정도로 하겠습니다. 이 숫자를 점점 100을 300으로 높이면 인식의 정확도가 65% 에서 70%로 올라가게 되고, nodes를 100에서 10000으로 높여보니까 정확도가 65%에서 95%로 향상되는 걸 확인해볼 수 있습니다. 그러니까 적은 숫자를 놓고 돌려봐서 정확도가 충분히 만족스러우면 그만해도 되고 아니면 이 숫자들을 점점 늘려가면서 정확도를 더 우리가 원하는 만큼 높여가면 됩니다. (화면의 코드를 보세요)


    We started with 784 input_nodes, 100 hidden_nodes, 10 output nodes.


    When we increased these numbers, accuracy rate has been improved from 65% to 70%. If we increase the number of nodes from 100 to 10000, we will see that the accuracy improves from 65% to 95%.


    So, when we run a few numbers, if we are satisfied with the given accuracy, we can stop. If not, we can increase the accuracy as much as we want by increasing these numbers.


    ---------------------------------------------------------------------

    # 입력, 은닉, 출력 노드의 수

    # Number of input, hidden, and output nodes

    input_nodes = 784

    hidden_nodes = 100   # 3000,  300 으로 높이면 65% 에서 70%

    # If you increase it to 3000, 300, it is from 65% to 70%

    output_nodes = 10    # 60000,  10000 으로 높이면 65% 에서 95%로 정확도 향상

    # Increased to 60000, 10000 to improve accuracy from 65% to 95%

    ---------------------------------------------------------------------

                                                             <---화면에 있으니 자막에는 없어도 됨)


    그 다음 학습률(learning_rate)을 0.3 으로 잡아주고, 그 다음에 신경망의 인스턴스를 생성한 후에, 학습 데이터 csv 파일을 불러옵니다. 불러올 수 있는 셋팅을 해주고, url을 다음과 같이 줘서 이 github에 있는 mnist_train_100을 불러옵니다. 그렇게 해서 신경망을 훈련 시킬 때, 이때 주기, 즉 epoch란 학습 데이터가 학습을 위해 사용되는 숫자인데, 일단 5로 시작해서 한번 확인해보겠습니다. 그래서 앞서 학습한 내용의 알고리즘으로 신경망을 훈련(training) 시키게 됩니다. (화면의 코드를 보세요)


    For the learning rate (learning_rate), it was given as 0.3 at first, and then it creates an instance a neural network and loads the training data csv file. When training a neural network in that way, the period, i,e. epoch, is the number that the training data is used for training. We practice it with 5 at first. So far, we have trained the neural network with the algorithm/code.


    ---------------------------------------------------------------------

    # 학습률 # Learning rate

    learning_rate = 0.3

    # 신경망의 인스턴스를 생성 # Create an instance of the neural network

    n = neuralNetwork(input_nodes, hidden_nodes, output_nodes, learning_rate)

    # mnist 학습 데이터인 csv 파일을 리스트로 불러오기

    # Load csv file, mnist training data, into a list

    import csv

    import urllib.request

    import codecs


    url = 'https://media.githubusercontent.com/media/freebz/Make-Your-Own-Neural-Network/master/mnist_dataset/mnist_train_100.csv'

    response = urllib.request.urlopen(url)

    training_data_file = csv.reader(codecs.iterdecode(response, 'utf-8'))

    training_data_list = list(training_data_file)

    response.close()


    # 신경망 학습시키기

    # 주기(epoch)란 학습 데이터가 학습을 위해 사용되는 횟수를 의미

    # Train neural network
    # Epoch means the number of times the training data is used for learning.

    epochs = 5


    for e in range(epochs):

        # 학습 데이터 모음 내의 모든 레코드 탐색

    # Explore all records within the training data set

        for record in training_data_list:

            all_values = record

            # 입력 값의 범위와 값 조정

    # Adjust the range and value of the input value

            inputs = (numpy.asfarray(all_values[1:]) / 255.0 * 0.99) + 0.01

            # 결과 값 생성 (실제 값인 0.99 외에는 모두 0.01)

    # Generate result value (all 0.01 except the actual value of 0.99)

            targets = numpy.zeros(output_nodes) + 0.01

            # all_values[0]은 이 레코드에 대한 결과 값

    # all_values[0] is the result value for this record

            targets[int(all_values[0])] = 0.99

            n.train(inputs, targets)

            pass

    pass

    ---------------------------------------------------------------------

                                                             <---화면에 있으니 자막에는 없어도 됨)


    데이터를 불러와서, 신경망을 구한 후, 잘 구했는지 테스트를 하고, 테스트 데이터 모음 내의 모든 레코드를 비교해서 정답이 어떤 것에 가장 가까운지 정답 또는 오답을 리스트에 추가하고, 정답인 경우 성적표에 1을 더하고 정답이 아닌 경우 성적표에 0을 더해서 정답의 비율인 성적을 계산해서 출력해줍니다. 그렇게 돌려보면 우리가 첫 이미지를 맞는 정답으로 찾았다는 비율이 대개 70% 정도가 된다고 결과가 나옵니다. (화면의 코드를 보세요)


    Now we load the data, find the neural network and test it. Compare all records in the data collection. If it is correct, we say 1, and if not, we say 0. The ratio of the correct answer can be found with this test. We could see that our ANN worked to give 70% of accuracy with the given conditions.

    ---------------------------------------------------------------------

    # mnist 테스트 데이터인 csv 파일을 리스트로 불러오기

    # Load csv file, mnist test data, into list

    url = 'https://media.githubusercontent.com/media/freebz/Make-Your-Own-Neural-Network/master/mnist_dataset/mnist_test_10.csv'

    response1 = urllib.request.urlopen(url)

    test_data_file = csv.reader(codecs.iterdecode(response1, 'utf-8'))

    test_data_list = list(test_data_file)

    response1.close()


    # 신경망 테스트하기 # Testing the neural network

    # 신경망의 성능의 지표가 되는 성적표를 아무 값도 가지지 않도록 초기화

    # Initialize the report card, which is an indicator of the performance of the neural network, to have no value

    scorecard = []


    # 테스트 데이터 모음 내의 모든 레코드 탐색

    # Explore all records within the test data set

    for record in test_data_list:

        all_values = record

        # 정답은 첫 번째 값

    # Correct answer is the first value

        correct_label = int(all_values[0])

        # 입력 값의 범위와 값 조정

    # Adjust the range and value of the input value

        inputs = (numpy.asfarray(all_values[1:]) / 255.0 * 0.99) + 0.01

        # 신경망에 질의

    # Query the neural network

        outputs = n.query(inputs)

        # 가장 높은 값의 인덱스는 레이블의 인덱스와 일치

    # The index of the highest value matches the index of the label

        label = numpy.argmax(outputs)

        # 정답 또는 오답을 리스트에 추가

    # Add correct or incorrect answer to the list

        if (label == correct_label):

            # 정답인 경우 성적표에 1을 더함

    # If the answer is correct, add 1 to the report card

            scorecard.append(1)

        else:

            # 정답이 아닌 경우 성적표에 0을 더함

    # If the answer is not correct, add 0 to the report card

            scorecard.append(0)

            pass

        pass


    # 정답의 비율인 성적을 계산해 출력

    # Calculate and print the grade, which is the percentage of correct answers

    scorecard_array = numpy.asarray(scorecard)

    print("performance = ", float(scorecard_array.sum()) / scorecard_array.size)

    ---------------------------------------------------------------------

                                                             <---화면에 있으니 자막에는 없어도 됨)


     이것이 의미하는 바는, (우리가 실습을 해서 얻은 결과는) 위의 알고리즘을 돌려보면, 학습 데이터를 100개, 테스트 데이터를 10개 사용했을 경우, 이 신경망의 정확도는 70%임을 확인할 수 있다는 것입니다. 나쁘지 않습니다. 시간도 별로 안 걸렸습니다. 그런데 이 학습 데이터의 수를 점점 늘려서 코드의 신경망을 좀 더 많은 약 60,000개의 학습 데이터를 갖고 학습시키고, 10,000개의 테스트 데이터에 적용을 하게 되면, 즉 학습 데이터를 100개에서 60,000개로 테스터 데이터를 10개에서 10,000개로 늘려 가면, 같은 알고리즘이 (사람이 구분하는 것보다 정확한) 대략 95%의 정확도를 갖는다는 것을 알게 됩니다.


    After we run the above algorithm (the result of our practice), we learned the following. When we started this ANN with 100 training data and 10 test data, the accuracy of this neural network was 70% which is not bad. It took only a very short time.


    However, as the number of data is increased as 60,000 training data and 10,000 test data from 100 training data and 10 test data, we will see that the accuracy of the same algorithm is increase to roughly 95% accurate (more accurate than humans does).


      MNIST 데이터 세트의 손 글씨 숫자를 인식한 사례는 신경망을 학습시키기 위한 데이터가 충분히 확보되었기에 가능했던 것입니다. 이전에는 사용할 수 있는 데이터가 많지 않았고 컴퓨터에 시간도 오래 걸렸지만, 요즘은, 수많은 Big Data 들이 매일 생산되고 있으며, 인공지능 시스템이 많은 양의 학습 데이터를 이용하여 점점 더 나은 추론을 하고 있습니다.


    The case of recognizing handwritten numbers in the MNIST data set was possible because enough data was available to train the neural network.


    Previously, there was not much data available and the computer took too much of time, but nowadays, a lot of Big Data is being produced every day, and AI systems has been improved getting better and better with this large amounts of training data.


      이 MNIST 숫자 인식 사례와 마찬가지로, 음성 또는 신호를 인식하거나 문자를 인식하는 것도 유사한 방법으로 가능합니다. 또 더 나아가서 손글씨나 숫자, 음성, 신호, 문자를 인식하는 것보다 더 높은 수준인, 우리의 말 즉 자연어(Natural Language)를 처리하는 연구는 우리가 일상생활에서 사용하는 언어를 컴퓨터가 의미를 분석하여 처리하도록 하는 것으로, 단순히 음성 또는 문자를 인식하는 것 이외에 말하는 사람이 의미한 것, 필요로 하는 것들을 추론을 하게 합니다.


    As with the example of [Recognition of Handwritten number with MNIST dataset], it is possible to recognize speech or signals or to recognize text in a similar way. Furthermore, a research that processes our speech, (that is, our natural language, which is more difficult than that of recognizing handwriting, numbers, voices, signals, and letters) is on going.


    With the good process of Natural Language, ANN will analyze the meaning of what we want and will try to give us a best possible answers.  Then we will be able to make our better decision quickly in a easier way.


    예를 들어, 사용자가 스마트폰을 들고 "피자 먹고 싶다!"라는 말했을 때, 인공지능 시스템은 “피자 먹고 싶다!"라는 문장을 인식하는 것에서 더 나아가, 그것을 텍스트로 바꿔서, 이 경우에 음식(food)으로 표기된 ‘피자’라는 용어의 사용에 주목하여, ‘조리법’ 같은 용어가 본문에 없음을 파악해서, ‘“아, 이 사람은 피자를 만들겠다는 것이 아니라 피자를 사먹고 싶다”라는 의미이다.’라고 진의를 이해하고, 말하는 사람이 피자집을 찾고 있다고 판단한 후, 검색을 시작합니다.


    For example, when I say "I want to pizza!" with my smartphone, the AI system goes further to understand the meaning. 'Oh, this person does not want to make pizza, but wants to buy or eat pizza from recognizing there is no term in the text such as 'recipe'. 'After AI understand the true meaning of what I said', AI understand that I said "I want to eat pizza!".


    다음으로 GPS 정보를 이용해서 말하는 사람의 주위에, 피자를 제공하는 식당을 찾아서 그 근접성과 가성비 등으로 순위를 매진 후, 이 사용자가 그 동안 사용한 피자 식당들의 기록이 있다면, 그 중 자주 가는 식당으로 인공지능 시스템이 추천을 해줍니다. 이런 방법으로 식당을 찾아서 식당의 이름과 주소를 파악한 후에, 메시지를 자연어 문장으로 생성한 후, 목소리로 다음과 같이 “여기서부터 1km 정도 떨어진 곳 오른쪽에 ***라는 피자집이 있습니다”라고 알려주는 것입니다. 음성인식은 물론 문자인식, 얼굴인식 등도 이런 방식의 인공지능 시스템을 활용하고 있다고 이해하시면 됩니다.


    Next, AI use GPS information to find restaurants that serve pizza near by me. And AI give ranks among them by the distance and cost-effectiveness. So AI makes recommendations.


    After finding a restaurant in this way and figuring out the name and address of the restaurant, the message is generated in natural language sentences. Will tell "There is a Gino's Pizza house in the right of 1km ahead."


     

    연습문제로는 숫자 인식, 얼굴인식, 음성인식, 자연어생성 기술이 현재 어느 정도 사용되고 있는지 알아보시고, 인공지능에 활용되는 수학에 대해서 알아보시고 공유해주시면 되겠습니다.


    As an exercise, you can try to find out [How AI has been used in face recognition, voice recognition, and Natural Language Generation in recent?]. And we may think and share the mathematics that are used in such AI system.


    이 사진은 이번 2020년 1월 CES에 전시된 자율주행 전기차의 내부인데, 엄청난 변화가 이루어져 있고 우리는 차 안에서 어디로 나를 모셔달라고 이야기만 하면 알아서 그 말을 인식하고 내비게이션으로 최적경로를 찾아서, 스스로 운전하며 안전하게 나를 데려다 줄것입니다. 그동안 나는 책을 읽고, 회의준비 하면서 도착해서 할 일을 준비할 것입니다. 이런 사회에 거의 다가와 있다는 것입니다.


    The photo shows the interior of a self-driving electric car on display at the CES in January 2020. There has been a huge change. Now if we ask AI to take me to Daejon KTX station, AI will recognize the meaning and find the best route through navigation and drive me safely to Daejon KTX station. In the meantime, we will read a book, prepare for the meeting, and prepare for the work to be done at there. It will be a life style that we will enjoy sooner or later.


    1차시에서는 여기까지 하고 2차시에서는 직접 이 내용을 실습하도록 하겠습니다. 수고하셨습니다.


    In the first lecture, we close up at this point and we will practice these code in the next lecture. Thank you.


    [14-2 pre]


    반갑습니다. 이제 마지막 시간이네요.


    Welcome to the final lecture of this class.


    지난 14주 동안 여러분들은 인공지능에 필요한 기본적인 다양한 수학내용들을 학습하셨습니다.


    Over the past 14 weeks, you have learned a variety of basic math contents required for artificial intelligence.


    이번 시간에는 ‘MNIST 데이터 셋을 활용한 손글씨 숫자 인식 사례’를 실습하고 그 동안 배워온 전체적인 내용들, ‘기초수학, 선형대수학, 미적분학, 통계학’의 주요내용들과 그것들을 활용해서 실제 ‘PCA’를 배우고 그것을 인공지능, 인공신경망에 직접 적용해본 전체 과정을 복습하면서 이번 학기를 마무리하도록 하겠습니다.


    In this class, we will practice the case on 'Handwritten Numbers Recognition Using MNIST Data Set', and review the whole contents learned we have so far, the main contents of 'Basic Mathematics, Linear Algebra, Calculus, Statistics', 'PCA', and Artificial Neural Network. I will finish this class by reviewing the entire process of learning and applying it directly to artificial intelligence and artificial neural networks.


    [14주차 2강] 안녕하세요. 14주차 두 번째, 마지막 강의입니다.


    [Lesson 2, Week 14]  Welcome~. Today we will have the last and final lecture of this class.


    지난 시간에 우리는 14주차 마지막 부분인 <MNIST 숫자인식> 알고리즘에 대해서 자세히 살펴보았습니다. 이번 시간에는 그 내용을 직접 실습하도록 하겠습니다.


    We just had a closer look on the <Number Recognition from the MNIST (Modified National Institute of Standards and Technology)  Dataset> in the first lecture of this week.


    Today we will practice that the <Number Recognition from the MNIST handwritten digit database> in http://matrix.skku.ac.kr/math4ai-intro/W14/.


    우리는 인공신경망을 활용한 손 글씨 숫자를 인식하는 사례를 github에 있는 데이터로부터 가져와서, 코드를 만들어서 구현했습니다. 그 내용은 미리 설명했듯이 library를 불러오고 csv 파일을 이 주소에서 불러와서, 이미지가 무엇인지를 파악했지요. 이것을 한번 실행해보도록 하겠습니다.


    We have revised a Python Code in  https://github.com/freebz/Make-Your-Own-Neural-Network to a Code in Sage to practice the <Number Recognition from the MNIST handwritten digit database in github>. This is a good example of using artificial neural networks. As we explained earlier, we call a library and get ‘a csv file’ from the page http://yann.lecun.com/exdb/mnist/ to figure out what the image was. Let's start.


    그러면, 이 데이터가 7이라는 숫자입니다. 이 데이터를 인식해서 ‘이 숫자는 7에 가장 가까운 숫자다’, ‘당신이 손으로 쓴 숫자가 아마 7일 것이다’라고 예측을 해준 것입니다. 이 training data가 7이었으므로 위의 이미지를 클릭하면 7이라는 숫자가 보여지고, 이 training data를 list[1]으로 바꿔주면 이 숫자가 다른 숫자로 나올 것입니다.


    We see this data is the number 7. The above code will recognize the data, and it will predicts that 'this number is more like to the number 7 in the data set', so 'AI can tell your hand-written number seems to be 7.' Since this first training data was 7, if we click [foo.png] after running ‘this text Code’, we will see the image of the number 7. And if we change this training data to the list [1], another number will come out.


    실제로 한번 바꿔서 실습을 해볼까요? 1을 한번 넣어서 실행을 해보면, 다른 숫자가 나올 것입니다. 그리고 이 숫자를 어떻게 인식하는 지, 이 숫자를 행렬로는 어떻게 인식하는지, 컬러이면 그 색깔을 행렬의 성분에 반영하여 표시했고, 이 행렬을 784×1 벡터로 표현해서 이 벡터가 들어갔을 때, 이 신경망이 이것을 행렬로 이해할 수가 있다고 했습니다.


    Let's change the number that we want to recognize and practice. If we put an other number in 'here' and run it, we'll get another image.


    Let's think about how ANN(artificial neural networks) recognizes this numbers. It recognize this number as a 28×28 matrix, (and if it is a photo, the darkness can be presented by the number of entries in a 28×28 matrix). This matrix will be considered as a 784×1 vector and the neural network starts to analyze it.


    그 벡터를 행렬에 곱해서 나온 output을 원래 예측하는 답과 비교해서 오차(error)를 계산한 후에 그것을 최소화하는 경사하강법(Gradient Descent) 알고리즘을 적용시키면서 훈련을 계속 시켜서 점점 더 정확한 값을 내놓는 신경망으로 만들어내는 과정을 학습했습니다.


    The vector is multiplied by a matrix (function) and is compared with the predicted output to compute the error. Afterwards, we learned the process of producing more and more accurate values by continuously applying the gradient descent algorithm that minimizes it (this is called the Backpropagation algorithm).


    이것이 우리가 가지고 있는 데이터들입니다. 데이터를 같이 봅시다. 이것이 테스트 데이터입니다. Neural Network을 만드는데 필요한 설명들이 소개되어 있는 내용이며, 여기서는 앞에서 설명했듯이, 신경망 클래스의 정의부터 시작해서 training 시키는 과정, 또 주기를 주고 노드의 수라든지 입력, 은닉, 출력 노드의 수를 조정하는 위치와 그리고 다른 알고리즘들은 그대로 활용하면 됩니다. 그래서 실제 아까 보신 신경망의 훈련의 알고리즘을 구현시켜보면, 다음과 같이 performance가 이번에는 0.5로 나왔습니다. 이는 50%의 정도 정확도를 가지고 있습니다. 이 숫자를 예를 들어서, 아까 보셨듯이 epochs를 더 늘려가고 데이터 숫자를 더 늘리면 더 정확해 집니다.


    These are the data we have. This code describes the process of training, the number of nodes, input, hidden layers, and output nodes. That can be adjusted according to our need.


    If we run this algorithm/code of training to a neural network. We have an output with a performance rate is 0.5: This means that it has an accuracy of 50%. As you saw earlier, we will have better accurate results by increasing <this number of epochs> and by increasing <the number of data>.

     

    주기를  7 정도로 늘려 볼까요? 그러면 performance가 10% 만큼 더 올라갔습니다. 데이터 수를 늘리거나 훈련시키는 주기를 늘려가거나 하면서 점점 신경망의 정확도를 높여갈 수 있습니다. 이것은 지금 간단한 개인 서버를 사용하는 거지만, 고성능 컴퓨터를 사용하면 이 데이터를 엄청, 지금부터 100배, 1000배 늘리더라도 지금 보셨듯이 바로 답을 확인할 수 있으니까, ‘우리가 원하는 99%, 98% 정도 정확도를 갖는 알고리즘을 만들기가 어렵지가 않다’라는 것을 이해할 수 있을 것입니다.


    By increasing this period to 7, we can improve the performance by 10%. We can increase the accuracy of the function (Matrix) by this neural network while increasing the number of data and training cycles (epoch).


    We used a personal server with a small number of data for this practice, but we can have a better answer right away when we increase the size of our data. It can be understood that it is not difficult to make a minor change on the above code to have an output with about 98% accuracy if we use a high-performance computer/server and big data.


    그리고 아까 보았듯이, 인공신경망을 훈련시켜 짜놓으면, 우리가 자연어처리 시스템을 이용해서 “아, 지금 어디를 가고 싶다.” 얘기하면 음성을 문자로 인식해서 ‘여기서 어떤 길로 가면 시간은 어느 정도 걸립니다.’ 하고 바로 답이 오게 되고, 또 “부산가는 KTX를 예약해 주세요.” 하면 알아보고 ‘지금 예약할 수 있는 자리는 몇 시에 어디에 있고 가격은 얼마입니다. 그리고 다음 기차는 몇 시에 출발하고 좌석은 어디 어디에 비어 각각의 가격은 얼마입니다.’ 이렇게 답을 주며, 그렇다면 두 번째 것으로 예약을 해달라고 얘기하면 인공지능이 대신 예약을 해 주는 것입니다. 우리가 배운 과정을 거쳐서, 여기에 자연어 처리까지 보태지면 이루어지는 일입니다.


    As we saw earlier, we can train an artificial neural network with Data, and set a system with a natural language processing system.


    Now we start. Let's talk to the natural language processing system, by saying "Hi! Siri, I want to go Busan tomorrow." When I say it, my voice is recognized as a text message, and the answer can be given as, '(1) 3 hours will take to get there!' And when I say, 'Please book a KTX ticket to Busan,' (after searching related informations instead of me, the following answer can be given '(2) There are tickets of seats in KTX that depart at 8AM and 1PM, and the price is US$50-, which one you like me to book for you? etc, including the choice of seats available.' If I ask for a reservation of the second choice, this AI secretary will make a reservation and type my credit card information to proceed with a payment.


    We can add more options in that process as we learned. These are the situations when the option of natural language processing is added on our ANN.



    이 과정에서 미리 얘기했듯이, 행렬 Ax=b라는 연립방정식을 푸는 여러 가지 해법들과 Gradient Descent Method, Singular Value Decomposition, Principle component analysis 등 지난 13주간 살펴 본 지식들이 모두 활용된 것입니다.


    As we mentioned before, all mathematical knowledge, including various theories on the system of linear equations (Ax=b), and the gradient descent method, the SVD, PCA, ANN that we have covered over the past 13 weeks, has been utilized to do such task.


    이제 14주 동안 Introductory Mathematics for Artificial Intelligence 내용을 간단히 복습해 보도록 하겠습니다. 최근에 인공지능과 대학수학 온라인 Zoom 심포지엄을 가졌습니다. 이 사진은 그 때 찍은 단체사진입니다.


    Let's take a quick review of the last 14-week lectures on Introductory Mathematics for AI. We recently held an online Symposium on <AI and University Mathematics>. This photo was a group photo taken at that time.


    우리는 이번 학기에 다음과 같이 인공지능에 필요한 기초수학으로 함수 그래프의 그래프를 그려서 축과 만나는 방정식의 해와 함수의 그래프들이 만나는 교점을 그리고 그 그래프를 확대, 축소해서, 구하거나 또는 근사식을 통해서 점점 더 가까이 가게 하면서 언제든지 다양한 다항식의 근을 구하는 방법을 배운 후, 2장 <인공지능과 행렬>에서는 데이터로부터 벡터, 행렬, 텐서로 정의했고, 그런 데이터들을 분류하는 방법, 크기로 분류하거나 방향으로 분류하거나하는 분류와, 선형연립방정식의 해집합, solution set을 구하는 방법, 가우스 소거법 등, 정사영과 최소제곱문제를 학습했고, LU-분해, QR-분해 특히 Singular Value Decomposition을 배웠으며, Singular Value Decomposition은 이후의 챕터에서 반복해서 계속 쓰였습니다.


    In this semester, we learned how to plot the graph of a function to  solve equations. An intersection where the graphs meet will be an answer.  We can zoom in that intersection point to get a reasonable numerical solution. We also learned how to get roots for various polynomials through numerical approximation. In Part 2 <Matrix and Data>, we defined Data as vector, matrix, and Tensor, we learned how to classify such Data in size or direction. The solution set of a system of linear equations, Gaussian elimination, the Projection and the Least Squares Problem were covered. LU-decomposition, QR-decomposition, especially SVD were introduced. And SVD has been used all over the remaining part of the book.


    3장 <인공지능과 최적해>에서는 미분, 적분 내용들이 있었는데, 그 중에서도 특히  최댓값과 최솟값, 극대 극솟값을 구하는 방법과 경사하강법 알고리즘을 이용해서 최소제곱문제의 해를 구하는 방법을 배웠고. 그리고 4장 <인공지능과 통계> 부분에서는 counting method, 확률변수, 이산확률분포, 연속확률분포, 조건부 확률, 대수의 법칙을 배웠으며, 분산, 공분산, 공분산 행렬을 학습했습니다.


    In Part 3 <AI and Optimal Solution>, we have covered the essence of <Calculus contents>. We have learned how to find the local maximum, local minimum, and Absolute maximum and Absolute minimum of a given function on an interval. And we learned how to solve the Least Square Problem using the <Gradient descent algorithm>. And in Chapter 4 <Artificial Intelligence and Statistics>, we have learned some of basic counting methods, random variables, discrete probability distribution, continuous probability distribution, conditional probability, the law of large number, variance, covariance, and covariance matrices.


    5장 <인공지능>에서는 공분산 행렬에 Singular Value Decomposition을 적용해서 주축을 찾는 Principal Component Analysis를 배웠습니다. 이를 통하여 원래 Big Data를 다룰 수 있는 비슷한 성질을, 좋은 성질을 가지고 있는 작은 사이즈의 행렬로 바꾸어서 계산함으로써 계산 시간을 엄청나게 줄이는 Rank reduction을 학습했으며, 그런 이론을 가지고 Artificial Neural Network을 디자인 하는 이론, 특히 오차 역전파법을 학습하였고, 그것들을 이용하여 인공신경망을 이용하여 실제 MNIST 데이터 set을 활용하여 숫자를 인식하는 이론과 실습을 학습하였습니다. 그리고 숫자 인식 인공신경망이 그대로 문자인식, 얼굴인식, 코드인식, bar 코드인식, 더 나아가서 음성인식 등 다양한 분야로 확장될 수 있다는 것을 배웠습니다.


    In Chapter 5 <Artificial Intelligence>, we learned the Principal Component Analysis to find the main axis by applying Singular Value Decomposition to the covariance matrix. Through this, we learned about a rank reduction which can reduce our computing time dramatically while it preserves the main properties of the original Data. While designing artificial neural networks, we learned the Backpropagation algorithm. With all of the above, we learned the theory/practice of recognizing numbers using an artificial neural network using the actual MNIST Data set. And we learned that the number recognition artificial neural network can be extended to various fields such as character recognition, face recognition, code recognition, bar-code recognition, and even voice recognition.


    이제, 이번 학기를 여기서 마치면서, 이번 주 14주 차에 학습한 내용을 한번 다시 돌아보겠습니다.


    Now, we will close out this semester, let's look back what we learned in this week.


    이번 14주차에는 MNIST 데이터 set을 활용한 숫자인식 실습을 했고, 우체국에서 우편번호를 인식하는 인공신경망이 학습하는 원리를 학습하고 실습했습니다. 그 내용은 숫자 데이터 set으로부터 실제 다른 사람이 편지봉투에 쓴 숫자를 이것과 비교해서 이 숫자가 어느 숫자에 가깝다는 것을 알려주는 그 알고리즘이 이 인공신경망 안의 알고리즘, 즉 오차 역전파법이라는 것을 자세히 학습하였습니다.


    In this Week 14, we learned and practiced the principle of learning by artificial neural networks that recognize postal codes at the post office. The content is covered in detail from the numerical data set, to the <Backpropagation algorithm> that really compare the number on the envelope with this one, to indicate that this number is close to a certain number.


    이제 14주 강의를 마치고 그 동안 학습한 내용들 또 토론한 내용들을 모아서 제출하는 과제를 샘플로 제공했으니 참고해서 보시기 바랍니다. 그 동안 수고 많았습니다.  그리고 14주 전체 과정을 같이 해주셔서 감사합니다. 축하드립니다.


    Now we finally finished all lecture for this semester. You may have a Final Homework to submit. The following can be a sample of it that we provide for you.

     

    Thank you for your hard work over the last 14-weeks for <Introductory Math for AI>.   You were wonderful and I enjoyed this class.

     

     Congratulations!!


    일반인을 위한 K-MOOC <인공지능을 위한 기초수학 입문>

                                             http://matrix.skku.ac.kr/math4ai-intro/

     

    Part 0.  인공지능에 필요한 기초수학 * 1주차

    Part 1.  인공지능과 행렬  * 2-6주차

    Part 2.  인공지능과 최적해 (미분) * 7-9주차 * Midterm PBL

    Part 3.  인공지능과 통계  * 10-11주차

    Part 4.  주성분 분석과 인공신경망  *12-14주차 * Final PBL


      <K-MOOC 인공지능을 위한 기초수학 입문>은 인공지능이 어떤 수학적 원리로 작동하는지를 이해하는데 필요한 기본적인 수학을 고등학교 1학년 정도의 수학 지식을 갖춘 일반인은 누구라도 쉽게 관련된 행렬, 도함수, 통계 내용을 이해하고 실습할 수 있도록 서술하였다.

      현재 인공지능은 우리가 느끼지 못하는 사이에 우리 삶의 거의 모든 곳에서 사용되고 있다. 우리는 인공지능에 대해 무작정 두려워하기보다 인공지능이 도대체 무엇인지 그리고 어떻게 작동되는지 기본 원리를 이해하면 된다. 이 책은 바로 그런 취지에서 준비되었다. 『인공지능을 위한 기초수학 입문』은 고등학생과 일반인을 대상으로 주 저자가 쓴 수학동아 『주니어매스』 원고와  대학생용 교재 『인공지능을 위한 기초수학』을 활용하여 미국, 일본, 중국의 고등학교 인공지능 수학 교수・학습 내용의 순서에 따라 우리나라 고등학교 2학년 이상이면 누구나 이해할 수 있도록 K-MOOC 교재와 강의록으로 완성하였다. 인공지능에 필요한 전반적인 용어와 개념을 접하고, 인공지능이 어떻게 작동하는지를 이해하는데 필수적인 수학 콘텐츠로 선형대수학, 다변수 미적분학, 기초 통계와 확률, SVD, 주성분 분석(PCA) 및 그래디언트 알고리즘 그리고 파이썬과 R 코드 및 실습실을 포함하는 내용을 담았다. 먼저, 인공지능의 알고리즘을 이해하는 데 필요한 최소한의 수학 개념을 고교 및 대학 저학년 수학 과정의 수준으로 설명한다. 이후, 앞서 배운 개념들이 실제로 인공지능을 개발할 때 어떻게 쓰이는지 잘 알려진 예와 알고리즘을 이용하여 구체적으로 설명한다. 특히 본 교재에는 수학이론 및 수학적 풀이와 관련되는 파이썬(Python) 및 R 코드를 제공하고, 클라우드 컴퓨팅 실습실에서 복잡한 계산과 데이터 처리 및 시각화를 직접 활용할 수 있도록 준비하였다.


    그림입니다.
원본 그림의 이름: KakaoTalk_20200528_094543939.png
원본 그림의 크기: 가로 1600pixel, 세로 900pixel


    Copyright @ 2020 SKKU Matrix Lab. All rights reserved.,          
    Made by
    Prof. Sang-Gu Lee (이상구) sglee at skku.edu

           그림입니다.       그림입니다.