Tom M. Mitchell,卡内基梅隆大学的教授,讲授机器学习等多门课程;美国人工智能协会(AAAL)的主席;美国 Machine Learning 杂志、国际机器学习年度会议(ICML)的创始人。
AI导读
核心看点
系统阐述概念学习、决策树、神经网络等核心算法
融合统计学、AI、信息论等多学科理论背景
深入解析归纳与分析学习结合及隐含假设
读者共识
机器学习领域经典之作,虽老但基础扎实
内容精炼无废话,但理论性强阅读门槛高
适合打基础,部分过时内容需结合新书补充
精彩摘录
"The inductive learning hypothesis. Any hypothesis found to approximate the target function well over a sufficiently large set of training examples will also approximate the target function well over other unobserved examples"
"We shall see that most current theory of machine learning rests on the crucial assumption that the distribution of training examples is identical to the distribution of test examples. Concept learning. Inferring a boolean-valued function from training examples of its input and output."
"As illustrated by these first two steps, positive training example may force the S boundary of the version space to become increasingly general. Negative training examples play the complimentary role of forcing the G boundary to become increasing specific."
"When gradient descent falls in a local minimum with respect to one of these weights, it will not necessarily be in a local minimum with respect to the other weights. In fact, the more weights in network, the more dimensions that might provide "escape routs" for gradient descent to fall away from the"
"The proof of this involoves showing that any function can e approximated by a inear combination of some samll region, and then showing that two layers of sigmoid units are sufficient to produce good local approximations."
"The only likely impact on the final error is that different error-minimization procedures may fall into different local minima. Bishop (1996) contains a general discussion of several parameter optimization methods for training networks A variety of methods have been proposed to dynamically grow or s"