Kevin P. Murphy is Associate Professor in the Department of Computer Science and in the Department of Statistics at the University of British Columbia.
"To understand these terms, you first need to understand the concept of likelihood. Assume you have a probability distribution - or rather family of such distributions - p(x;w) which assigns a probability to each data point x, given a specific setting of its parameters w. That is, different values of"
"In particular, we define machine learning as a set of methods that can automatically detect patterns in data, and then use the uncovered patterns to predict future data, or to perform other kinds of decision making under uncertainty"
"a property known as the long tail, which means that a few things (e.g., words) are very common, but most things are quite rare. Machine learning is usually divided into two main types. In the predictive or supervised learning approach. most methods assume that yi is a categorical or nominal variable"
"In our notation, we make explicit that the probability is conditional on the test input x, as well as the training set D, by putting these terms on the right hand side of the conditioning bar |. When choosing between different models, we will make this assumption explicit by writing p(y|x,D,M), wher"
"Regression is just like classification except the response variable is continuous."
"Instead, we will formalize our task as one of density estimation, that is, we want to build models of the form p(xi|θ). There are two differences from the supervised case: First, we have written p(xi|θ) instead of p(yi|xi, θ); that is, supervised learning is conditional density estimation, whereas u"
"Picking a model of the “right” complexity is called model selection, and will be discussed in detail below. zi is an example of a hidden or latent variable, since it is never observed in the training set."
"There are many ways to define such models, but the most important distinction is this: does the model have a fixed number of parameters, or does the number of parameters grow with the amount of training data? The former is called a parametric model, and the latter is called a nonparametric model. Pa"