Peter Harrington holds Bachelors and Masters Degrees in Electrical Engineering. He worked for Intel Corporation for seven years in California and China. Peter holds five US patents and his work has been published in three academic journals. He is currently the chief scientist for Zillabyte Inc. Peter spends his free time competing in programming competitions, and building 3D printers.
AI导读
核心看点
以Python代码实现经典机器学习算法
涵盖分类、回归及无监督学习核心内容
强调从底层构建代码理解算法原理
读者共识
优秀的机器学习入门实战指南
代码实现有助于直观理解概念
数学理论较浅需配合其他教材
精彩摘录
"Pros: High accuracy, insensitive to outliers, no assumptions about data Cons: Computationally expensive, requires a lot of memory Works with: Numeric values, nominal values The first machine-learning algorithm we’ll look at is k-Nearest Neighbors (kNN). It works like this: we have an existing set of"
"Pros: Computationally cheap to use, easy for humans to understand learned results, missing values OK, can deal with irrelevant features Cons: Prone to overfitting Works with: Numeric values, nominal values"
"General approach to decision trees 1. Collect: Any method. 2. Prepare: This tree-building algorithm works only on nominal values, so any continuous values will need to be quantized. 3. Analyze: Any method. You should visually inspect the tree after it is built. 4. Train: Construct a tree data struct"
"Logistic regression Pros: Computationally inexpensive, easy to implement, knowledge representation easy to interpret Cons: Prone to underfitting, may have low accuracy Works with: Numeric values, nominal values"
"The clear syntax of Python has earned it the name executable pseudo-code."
"With Python, you can program in any style you’re familiar with: object-oriented, procedural, functional, and so on."
"With Python it’s easy to process and manipulate text"
"Python is popular in the scientific and financial communities as well.A number of scientific libraries such as SciPy and NumPy allow you to do vector and matrix operations."
目录
Part 1: Classification
1 Machine learning basics
2 Classifying with k-nearest neighbors
3 Splitting datasets one feature at a time: decision trees
4 Classifying with probability distributions: Na�ve Bayes