← Back to Part 1: Quickly revise ML concepts

Machine Learning · Self-test

ML Concepts Quiz — Part 1

101 multi-select questions covering everything in Part 1, from what ML even is through gradient descent and probability. Select any options you think are true, then check — each question reveals its own worked explanation.

0/101Answered
0Correct

Section 1 · 8 questions

What Is Machine Learning?

Definitions, why rule-based systems break down, and where ML actually earns its keep.

Q1What Is Machine Learning?

Tom Mitchell's 1998 definition says a program learns from experience E with respect to task T and performance measure P if its performance at T, measured by P, improves with E. For a spam filter, which of the following are correctly matched to E, T, or P?

Select all statements you believe are true, then check your answer.

Explanation

The <P,T,E> triple is the standard way interviewers check you actually understand what 'learning' means operationally, not just as a buzzword. E is the data/interaction the system is exposed to (labeled emails). T is the task being performed at deployment time (classify a new email). P is a measurable score of how well T is being done (accuracy, F1, etc.) — it must be something you can compute, so 'total emails ever sent' isn't a performance measure of this system at all, and 'algorithm used to store emails' is an implementation detail, not the task.

Interview angle: when asked to design an ML system, stating the <P,T,E> triple explicitly (and picking a P that's actually aligned with business goals, e.g. weighted for false-positive cost) is a strong opening move in a system-design round.

Q2What Is Machine Learning?

Which statements correctly distinguish Arthur Samuel's (1959) definition of ML from Tom Mitchell's (1998) definition?

Select all statements you believe are true, then check your answer.

Explanation

Samuel's phrase ('the ability to learn without being explicitly programmed') is a great intuition-builder but isn't falsifiable/testable on its own. Mitchell's <P,T,E> formulation is what lets you write down, for any proposed system, exactly what experiment would show it 'learned.' Neither definition ties ML to a specific algorithm family (gradient descent, neural nets, decision trees, etc.) — that's a category error interviewers listen for.

Q3What Is Machine Learning?

A rule-based (hand-coded) system tries to recognize handwritten digits by measuring shape features like the number of closed loops. Based on the slide's OCR/digit examples, which statements are true?

Select all statements you believe are true, then check your answer.

Explanation

This is the core 'why ML' argument from the deck: for problems with huge natural variability (handwriting, animal appearance, language), enumerating explicit rules doesn't scale — you'd need an unbounded rule set to cover every case, and rules interact badly. The OCR example specifically shows that local shape matching without context produces plausible-looking but wrong output ('fooEball'), which is exactly why context-aware, learned models (even simple n-gram language models) outperform naive character-shape classifiers. No amount of manually added rules asymptotically beats a system that can generalize from data, because the rule-writer runs out of foresight before the data runs out of edge cases.

Q4What Is Machine Learning?

According to the 'When is ML useful?' criteria discussed, which conditions should hold before reaching for a learned model instead of hand-written logic?

Select all statements you believe are true, then check your answer.

Explanation

This is a practical gate every ML engineer should run before proposing a model: if there's no real pattern, you'll just fit noise; if the pattern is trivially expressible as a formula (e.g. converting Celsius to Fahrenheit), a model is overkill and adds needless risk/latency/maintenance; and without enough representative data, even the right model architecture will fail to generalize. Class balance and linearity are neither necessary conditions — plenty of successful ML systems handle imbalance (with reweighting/resampling) and nonlinearity (via nonlinear models), so they don't gate the initial 'should we even use ML' decision.

Interview angle: this is the mental checklist behind 'when would you NOT use a neural network' — a classic system-design probe.

Q5What Is Machine Learning?

Which of the following correctly describe the relationship between AI, ML, and Deep Learning as nested fields?

Select all statements you believe are true, then check your answer.

Explanation

The classic nested-circles picture (AI ⊃ ML ⊃ DL) is a favorite whiteboard question. AI is the broadest field: any system that exhibits behavior we'd call 'intelligent,' including purely symbolic/rule-based planners and search algorithms with zero learning (e.g. A* pathfinding, minimax game trees) — so 'every AI system is ML' is false. ML narrows this to systems that improve via data/experience. DL further narrows ML to models composed of many layers of learned, hierarchical representations (CNNs, transformers, etc.).

Q6What Is Machine Learning?

Which of these are cited as reasons for ML's recent surge in popularity, distinct from the core mathematical ideas (which are decades old)?

Select all statements you believe are true, then check your answer.

Explanation

A key interview distinction: the underlying theory (perceptrons 1957/58, backprop formalized in the 1980s, SVMs in the 1990s) is old. What changed circa 2010s was the surrounding ecosystem — GPU throughput for matrix math, elastic cloud compute, internet-scale datasets, and frameworks (PyTorch/TensorFlow) that made iterating on architectures fast for non-systems-experts. Knowing this story shows you understand ML's history isn't 'suddenly discovered math,' it's 'old math finally had the substrate to run at scale.'

Q7What Is Machine Learning?

Traditional programming and machine learning are often contrasted as two different 'functions' a computer computes. Which statements correctly describe this contrast?

Select all statements you believe are true, then check your answer.

Explanation

This flip is the single clearest way to explain ML to a non-technical stakeholder: instead of a human writing the transformation rules (the program) by hand, the computer infers the transformation from examples of correct input→output pairs, and the output of that training process (weights, split thresholds, coefficients, etc.) becomes the new 'program' used at inference time. At inference/serving time, you absolutely still need input data — 'no data needed' is wrong; only training-time labels become unnecessary at inference.

Q8What Is Machine Learning?

Which of the following are legitimate reasons an organization chooses ML-based automation over a human-in-the-loop decision process, per the 'Why ML?' motivations discussed (automation, speed, observer independence)?

Select all statements you believe are true, then check your answer.

Explanation

'Observer independence' means two different human raters (or the same rater on different days) often disagree — a calibrated model applies the exact same function every time, which is valuable in domains like radiology triage or content moderation where consistency itself has value. Speed/scale is the throughput argument. Critically, automating a decision does not remove the need for ongoing evaluation — in fact it raises the stakes, since silent model drift can now affect every decision at scale, so monitoring becomes more important, not less.

Section 2 · 9 questions

Supervised, Unsupervised & Reinforcement Learning

Telling the paradigms apart by what experience the learner gets, not by the algorithm.

Q9Supervised, Unsupervised & Reinforcement Learning

Which of the following are the three paradigms of machine learning explicitly covered in the slides?

Select all statements you believe are true, then check your answer.

Explanation

Supervised, unsupervised and reinforcement learning are the three foundational paradigms defined by what kind of experience/feedback the learner receives: labeled (data, label) pairs, unlabeled data only, or a sequence of actions with delayed rewards/penalties from an environment. Transfer learning and federated learning are real, important techniques, but they're orthogonal concerns (how you reuse a model, or how you train across decentralized data) that can apply within any of the three paradigms — they aren't a fourth co-equal paradigm in this taxonomy.

Q10Supervised, Unsupervised & Reinforcement Learning

In supervised learning, which statements about the 'learner' vs. the 'reasoner' are correct?

Select all statements you believe are true, then check your answer.

Explanation

This learner→model→reasoner pipeline is worth internalizing precisely because it maps directly onto MLOps terminology: 'learner' = training job, 'model' = serialized artifact/checkpoint, 'reasoner' = inference/serving. At inference time the true label is exactly what's unknown — if you had it, there'd be nothing to predict; that's the whole point of a held-out test set: labels exist for evaluation, but the model never sees them as input. Background/prior knowledge is optional — many supervised learners (e.g. plain linear regression from a CSV) work with zero domain knowledge injected.

Q11Supervised, Unsupervised & Reinforcement Learning

Which of these are correctly classified as classification tasks vs. regression tasks in supervised learning?

Select all statements you believe are true, then check your answer.

Explanation

The discriminator is the type of the target/label: categorical (a finite set of classes, possibly just 2 as in binary classification) → classification; real-valued/continuous → regression. A continuous 2D coordinate is regression (in fact multi-output regression), not classification, even though it 'looks' like it could be bucketed — unless you deliberately discretize the coordinate space into bins, at which point it becomes a classification problem by construction.

Q12Supervised, Unsupervised & Reinforcement Learning

An unsupervised clustering algorithm is given thousands of unlabeled customer-purchase vectors and must group similar customers. Which statements are true about this setting?

Select all statements you believe are true, then check your answer.

Explanation

Clustering is unsupervised because no example carries a 'correct group id' during training — the algorithm discovers structure (cohesive groups of nearby points) rather than matching pre-assigned labels; that's fundamentally different from classification, where a known label set drives training. Cluster quality absolutely can be evaluated without ground truth using internal metrics like silhouette score or within-cluster sum of squares, even though external metrics (e.g. adjusted Rand index against known labels) are also an option when labels happen to exist for validation purposes only.

Q13Supervised, Unsupervised & Reinforcement Learning

Which of the following are legitimate example applications of unsupervised learning mentioned in the deck (clustering, dimensionality reduction, association rule mining, generative models)?

Select all statements you believe are true, then check your answer.

Explanation

All four listed use-cases genuinely require no labels: clustering groups similar things, dimensionality reduction (e.g. PCA) finds a compact representation preserving structure, and association rule mining (e.g. market-basket analysis, 'customers who bought X also bought Y') finds co-occurrence patterns. Predicting a specific future numeric value from historical numeric values is textbook supervised regression — every historical (month, revenue) pair is itself a labeled training example, since 'last month's revenue' plays the role of a known target during training.

Q14Supervised, Unsupervised & Reinforcement Learning

You are teaching a child to ride a bicycle. According to the reinforcement-learning framing used in the slides, which statements are correct?

Select all statements you believe are true, then check your answer.

Explanation

Bike-riding is the deck's canonical intuition-builder for RL: there's no dataset of (situation, correct-handlebar-angle) pairs handed to the child up front (ruling out pure supervised learning), and it's not just 'find structure in static data' either (ruling out pure unsupervised learning) — instead, the learner tries actions, observes consequences (reward/penalty) from an environment, and adjusts a policy over a sequence of decisions, not a single-shot prediction. That sequential, feedback-driven, action-selection framing is exactly the Agent–Environment–State–Action–Reward loop formalized as a Markov Decision Process.

Q15Supervised, Unsupervised & Reinforcement Learning

In the standard reinforcement-learning model (Agent, Environment, State S_t, Action A_t, Reward R_t), which statements are correct?

Select all statements you believe are true, then check your answer.

Explanation

This is the standard RL loop: agent observes S_t, picks A_t, environment transitions and emits R_(t+1) and S_(t+1), and the loop repeats. The environment is not optional — it's the thing that actually defines the dynamics and reward function the agent is trying to learn about; without it there's nothing to interact with. And RL agents are explicitly optimizing for long-run reward (often formalized via a discounted sum), which is why a good RL policy sometimes accepts a worse immediate reward for a much better future payoff (e.g. sacrificing a chess piece for positional advantage).

Q16Supervised, Unsupervised & Reinforcement Learning

Which of these are examples of 'other' learning paradigms mentioned (weakly-supervised, semi-supervised, self-supervised) beyond the three main ones?

Select all statements you believe are true, then check your answer.

Explanation

These sit on the spectrum between fully supervised and fully unsupervised. Semi-supervised uses a small labeled + large unlabeled pool (common when labeling is expensive, e.g. medical imaging). Self-supervised constructs 'free' labels from the raw data's own structure — masked-language-model pretraining (BERT), next-token prediction (GPT), and image rotation/colorization pretext tasks are all self-supervised and span both text and vision (and audio, video, etc.). Weakly-supervised uses labels that are cheap but imperfect — e.g. labeling every pixel in an image with the image-level tag rather than doing precise pixel-level annotation.

Q17Supervised, Unsupervised & Reinforcement Learning

Which of the following correctly distinguishes a classification model's target from a regression model's target, using the price-of-house and cat-vs-dog examples from the slides?

Select all statements you believe are true, then check your answer.

Explanation

The algorithm family (classification vs regression) is a property of the target/label, not the input features — the same feature vector (say, square footage + location + age of a house) could feed a regression model predicting price or a classification model predicting 'sells in under 30 days: yes/no.' Classification supports any number of classes ≥ 2 (multiclass, e.g. digit recognition 0–9), and regression targets can be negative (e.g. predicting temperature change, profit/loss) — 'always positive' is a dataset property, not a definitional requirement.

Section 3 · 9 questions

Inductive Learning, Hypothesis Spaces & Bias

Why learning at all requires an assumption, and what kind of assumption an algorithm is making.

Q18Inductive Learning, Hypothesis Spaces & Bias

In inductive learning for the supervised setting, we assume there exists a true target function f(X) and we try to learn a hypothesis h that approximates it. Which statements are correct?

Select all statements you believe are true, then check your answer.

Explanation

This is the precise vocabulary interviewers expect: f is the unknown, unobservable ground-truth mapping; h is our approximation, drawn from a hypothesis space H (e.g. 'all straight lines,' 'all depth-3 decision trees'). Crucially, induction never guarantees h = f — inductive learning is fundamentally an ill-posed problem (more below) precisely because a finite training sample is consistent with infinitely many hypotheses, so we can only claim h approximates f well on the distribution the data came from, never that we've recovered f exactly.

Q19Inductive Learning, Hypothesis Spaces & Bias

Why is inductive learning described as an 'ill-posed problem'?

Select all statements you believe are true, then check your answer.

Explanation

'Ill-posed' here has a specific technical meaning: a well-posed problem has a unique solution determined by the given information; here, the given information (a finite labeled sample) is compatible with infinitely many candidate functions that all agree on the training points but diverge everywhere else. This is the formal justification for why every learning algorithm needs an inductive bias — some built-in preference (e.g. 'prefer simpler/smoother functions') that breaks the tie among all the hypotheses that fit the training data equally well.

Q20Inductive Learning, Hypothesis Spaces & Bias

Consider a binary classification problem with N training points in a 2D feature space. Without any restriction on the hypothesis, how many distinct labelings (functions from points to {positive, negative}) are theoretically possible, and what does this imply?

Select all statements you believe are true, then check your answer.

Explanation

Each of the N points can independently be labeled one of 2 ways, giving 2^N possible label assignments — with even 30 points that's over a billion, so brute-force search over 'all conceivable functions' is a non-starter. This is precisely the practical motivation for restricting the hypothesis space to something tractable (straight lines, decision trees of bounded depth, neural nets of fixed architecture) — that restriction, plus any preference among the remaining hypotheses, together constitute the model's inductive bias.

Q21Inductive Learning, Hypothesis Spaces & Bias

Which of the following correctly describe 'restriction bias' (a.k.a. absolute bias) as distinguished from 'preference bias' (a.k.a. relative bias)?

Select all statements you believe are true, then check your answer.

Explanation

This distinction shows up constantly in real system design: choosing a model class (linear model vs. gradient-boosted trees vs. deep net) sets your restriction bias — it hard-limits what functions are even reachable. Regularization, priors, or early stopping don't remove any hypothesis from being reachable in principle, they just make the optimizer prefer 'simpler' members of that same space — that's preference bias. Knowing which lever you're pulling (restriction vs. preference) is exactly how you'd explain, in an interview, why increasing model capacity vs. increasing regularization strength have different effects on the bias-variance tradeoff.

Q22Inductive Learning, Hypothesis Spaces & Bias

Ockham's razor is invoked as a heuristic in model selection. Which statements about it are correct?

Select all statements you believe are true, then check your answer.

Explanation

Ockham's razor is a heuristic/assumption, not a theorem — the deck is explicit that 'this is just an assumption and does not always work.' It's useful precisely because it's one reasonable inductive bias among many, and the No Free Lunch theorem tells us that no bias is universally best across every possible problem — you're always betting that your assumption matches the world you're actually operating in. 'Simpler' is also not strictly synonymous with 'fewer parameters' — e.g. a smooth, low-curvature function can be 'simpler' in a meaningful sense even with many parameters, if those parameters are constrained (regularized) to behave smoothly.

Q23Inductive Learning, Hypothesis Spaces & Bias

You fit a regression curve to 5 training points twice: once assuming the underlying relationship is 'current vs. voltage' (physically near-linear, Ohm's law) and once assuming it's 'day vs. stock price' (noisy and volatile). Which statements correctly reflect the lesson from this example?

Select all statements you believe are true, then check your answer.

Explanation

This is the deck's most vivid illustration of inductive bias in action: identical (x, y) coordinate values call for different curve shapes depending on domain knowledge about the data-generating process. Ohm's law is genuinely near-linear physics, so a straight line with the odd measurement-noise outlier is the right story. Daily stock prices are driven by many volatile, hard-to-model factors, so a smooth low-degree curve would actually underfit the real dynamics — a rougher fit may better reflect reality. The lesson: 'best fit' is never data-only, it's data + assumptions about the domain.

Q24Inductive Learning, Hypothesis Spaces & Bias

A car's speed-over-time plot could be drawn as either (a) a perfectly flat horizontal line or (b) a gently wavy curve with small ups and downs. Which statements correctly reflect the reasoning given in the slides?

Select all statements you believe are true, then check your answer.

Explanation

This example is used to build intuition for a somewhat counter-intuitive equivalence: 'simpler model' = 'more restrictive assumptions' = 'higher bias.' A flat-line speed model is simple to write down, but only matches reality if you assume away all the real-world complexity (traffic, signals) — that's a strong, restrictive assumption, i.e. high bias. The wavy curve requires fewer such assumptions (it lets speed vary freely to match what actually happens) and is therefore lower-bias but more complex. This is exactly backwards from the naive intuition that 'more assumptions = more accurate,' which is why it's worth drilling.

Q29Inductive Learning, Hypothesis Spaces & Bias

Which of the following are correctly described as inductive-bias decisions an ML engineer explicitly makes when designing a model, as opposed to something the data determines automatically?

Select all statements you believe are true, then check your answer.

Explanation

Everything in this list except the raw data values is a design choice that encodes an assumption about the world — model family (restriction bias), tree depth (restriction bias, controls the size of the reachable hypothesis space), L2 penalty (preference bias, favors small-weight solutions), and even architectural choices like convolution (restriction+preference bias baked in: 'nearby pixels matter more, and the same filter should apply everywhere in the image') are all inductive biases chosen by the practitioner. The raw data values themselves are just observations — they're what the bias gets applied to, not a form of inductive bias.

Q88Inductive Learning, Hypothesis Spaces & Bias

An interviewer asks: 'If I give you a training set and no other information, can you tell me the single best model to use?' Which responses correctly apply the ideas of inductive bias and hypothesis space?

Select all statements you believe are true, then check your answer.

Explanation

This is a great meta-question because it directly tests whether the No-Free-Lunch / inductive-bias lesson actually sank in versus being memorized as trivia. The 'correct' senior-level answer isn't a specific algorithm name — it's recognizing that the question itself is underspecified, and demonstrating the diagnostic questions (what's the target type, how much data, what's known about the domain's structure, what's the cost of different error types) that a working ML scientist would ask before committing to a hypothesis space at all.

Section 4 · 5 questions

Features & Feature Space

Turning raw data into the coordinates an algorithm can actually draw a boundary in.

Q30Features & Feature Space

Which statements correctly describe the relationship between raw data, features, and feature vectors, as introduced in the slides?

Select all statements you believe are true, then check your answer.

Explanation

Feature vs. feature vector is basic vocabulary: individual measurable properties (height, sepal width, pixel intensity) are features; stacking them for one example gives a feature vector, e.g. (sepal_length=5.1, sepal_width=3.5, petal_length=1.4, petal_width=0.2). Modern deep nets (CNNs on raw pixels, transformers on raw tokens) can skip hand-engineered features and learn representations end-to-end — but that's specifically a deep-learning capability; classical algorithms like logistic regression, SVMs with hand-picked kernels, or gradient-boosted trees on tabular data still typically need engineered/selected features to perform well.

Q31Features & Feature Space

Where can features come from, according to the slides ('How to Obtain Features')?

Select all statements you believe are true, then check your answer.

Explanation

Feature acquisition spans a spectrum from fully manual (a doctor recording a patient's blood pressure) to fully learned (a CNN's early layers learning edge/texture detectors on its own from raw pixels). Classical computer vision relied on hand-designed descriptors (SIFT, HOG, texture filters) before deep learning made raw-pixel-in, learned-features-out common. Features certainly don't require a physical instrument — a categorical flag, a ratio computed from two other columns, or a one-hot encoded string are all valid engineered features.

Q32Features & Feature Space

In a 2D feature space where x1 = height and x2 = weight are used to separate cats from dogs, which statements are correct?

Select all statements you believe are true, then check your answer.

Explanation

Feature space is literally the coordinate system the model reasons in — the choice of which features to include changes the geometry of the problem, sometimes making an otherwise inseparable problem separable (this is the core motivation behind kernel methods and feature engineering generally: project into a space where classes become linearly separable). But more features is not free — it increases the dimensionality of the space the model must learn to reason about, which (per the curse of dimensionality) typically requires more, not less, data to reliably estimate a good decision boundary, all else equal.

Q33Features & Feature Space

Which of the following correctly describes 'the data' X as discussed for building a supervised model — considering that X can be raw data or a feature vector?

Select all statements you believe are true, then check your answer.

Explanation

X (inputs) and Y (labels/targets) are always kept conceptually and practically separate — Y is what the model is trying to predict, never an input feature (leaking Y into X, even indirectly through a derived feature, is the classic 'target leakage' bug that silently inflates offline metrics and then fails in production). The raw-vs-engineered-features choice is a genuine engineering tradeoff: raw pixel/audio input pushes representation learning onto the model (which then needs more capacity and usually more data), while hand-engineered features front-load domain knowledge but require that expertise to be available and correct.

Q90Features & Feature Space

A recruiter dataset has a 'years of experience' feature ranging from 0–40 and a 'has a master's degree' binary feature (0/1). Which statements about combining these into one feature vector are correct?

Select all statements you believe are true, then check your answer.

Explanation

This is a very concrete, hiring-relevant version of the hemoglobin/creatinine scaling problem: in raw Euclidean distance, a 5-year experience gap contributes 25 to the squared distance while the entire binary degree feature can contribute at most 1 — meaning the degree feature is almost invisible to a k-NN/k-means model unless features are scaled first. Tree splits ('experience > 7.5?', 'has_masters == 1?') are evaluated per-feature independently of other features' units, so trees genuinely don't need this kind of scaling — this is one of the most commonly asked 'why don't we need to scale features for random forest/XGBoost' interview questions.

Section 5 · 16 questions

Bias, Variance & Generalization

Underfitting vs. overfitting, Ockham's razor, and why you can't drive both errors to zero at once.

Q34Bias, Variance & Generalization

Two models are fit to the same training data: Model 1 is a straight line (green), Model 2 is a high-degree polynomial that passes through every single training point exactly (red, wiggly). Which statements about their expected behavior are correct?

Select all statements you believe are true, then check your answer.

Explanation

This is the deck's core empirical demonstration of bias-variance tradeoff using repeated test sets. The wiggly high-degree curve minimizes training error (possibly to zero) by chasing noise unique to that particular training sample — that noise-fitting is exactly what makes its performance swing wildly (high variance) across different, independently drawn test sets. The straight line under-fits some of the fine structure (higher bias) but that same rigidity makes it robust/consistent across resamples (low variance). Crucially, lower training error does NOT imply lower test error — that's the overfitting trap.

Q35Bias, Variance & Generalization

Formally, which statements correctly define 'bias' and 'variance' of a model in the bias-variance decomposition sense?

Select all statements you believe are true, then check your answer.

Explanation

These are the textbook statistical definitions used throughout ML theory and directly quoted from the slides. Bias is a property of the model's average behavior relative to ground truth (a systematic error), while variance is a property of how sensitive that behavior is to which particular training sample you happened to draw — it is not about label noise itself, which is irreducible ('Bayes error'), separate from both bias and variance. In practice, bias/variance can indeed vary locally — e.g. a linear model might fit a linear region of the true function well (low local bias) but badly miss a curved region (high local bias there).

Q36Bias, Variance & Generalization

Which statements correctly connect model complexity to bias and variance?

Select all statements you believe are true, then check your answer.

Explanation

The bias-variance tradeoff is called a tradeoff precisely because these two error sources move in opposite directions as you slide the complexity dial — you cannot drive both to zero simultaneously with a fixed amount of training data. Underfitting = the model is too rigid to capture real structure (high bias) but is stable across resamples (low variance). Overfitting = the model is flexible enough to fit noise (low training-set bias) but that flexibility makes it wildly sensitive to which noise it happened to see (high variance). The practical goal is finding the complexity sweet spot, not maximizing complexity.

Q37Bias, Variance & Generalization

In the classic 'Error vs. Model Complexity' chart (training error monotonically decreasing, test error U-shaped), which statements are correct?

Select all statements you believe are true, then check your answer.

Explanation

This U-shaped test-error curve is one of the single most important charts in applied ML — it's what justifies practices like early stopping, cross-validated hyperparameter search, and held-out validation sets in the first place. The optimum sits where test/validation error bottoms out, which is essentially never the same point where training error is lowest (training error keeps improving, or at least doesn't get worse, all the way to the most complex/overfit end). Watching the train-test gap widen as you keep adding complexity is literally how engineers detect overfitting in a training log.

Q38Bias, Variance & Generalization

A team has only 5 training examples, all fairly similar to each other. Which statements correctly capture the risk discussed regarding small/homogeneous datasets and overfitting?

Select all statements you believe are true, then check your answer.

Explanation

This is a practical, often-overlooked corollary of the bias-variance discussion: a suspiciously perfect training fit is much easier — and much less meaningful — when you have few or very similar samples, because there's less true signal-vs-noise diversity to distinguish. As dataset size and diversity grow, a model that can still hit every training point exactly is almost certainly memorizing rather than generalizing (which is exactly why held-out validation, not training accuracy, is the metric that matters). This is also the core justification for data augmentation and collecting more diverse examples as an anti-overfitting technique, distinct from regularization.

Q39Bias, Variance & Generalization

Which of the following are practical techniques an ML engineer would use specifically to reduce overfitting (high variance) rather than to reduce underfitting (high bias)?

Select all statements you believe are true, then check your answer.

Explanation

All four listed techniques directly reduce variance: regularization constrains the effective hypothesis space toward 'simpler' solutions, more diverse data makes it harder to memorize noise, reducing capacity shrinks the reachable hypothesis space, and early stopping halts training before the model starts fitting noise past the point where validation loss starts rising. Increasing capacity does the opposite — it's the standard fix for underfitting (high bias), not overfitting, and would typically make an already-overfit model worse.

Q40Bias, Variance & Generalization

Which of the following are practical signs, observed during model training/evaluation, that a model is underfitting rather than overfitting?

Select all statements you believe are true, then check your answer.

Explanation

Underfitting shows up as consistently poor performance everywhere (train and validation both bad, roughly equal), because the model's hypothesis space is simply too restrictive to capture the pattern regardless of how much data you throw at it — which is exactly why 'just collect more data' doesn't fix a genuinely underfit model, you need more capacity/features/a richer model family instead. A large, growing train-validation gap is the signature of overfitting, not underfitting — that's the opposite pattern.

Q41Bias, Variance & Generalization

A model shows training accuracy = 99% and validation accuracy = 68%. Which are appropriate next diagnostic/remedial steps?

Select all statements you believe are true, then check your answer.

Explanation

A 31-point train-validation gap is a textbook overfitting signature, so the standard playbook is: (1) regularize/simplify, (2) get more/more-diverse data, and (3) rule out leakage — a subtly leaked feature or overlapping train/val rows can produce exactly this pattern (suspiciously high training performance) and is a very common real-world root cause that's easy to overlook if you jump straight to 'add regularization.' Increasing model complexity would almost certainly make this specific problem worse, not better — that's the fix for the opposite symptom (underfitting).

Q42Bias, Variance & Generalization

Which statements correctly summarize the 'sweet spot' goal in the bias-variance tradeoff, per the classic 3-panel picture (high variance / high bias / good balance)?

Select all statements you believe are true, then check your answer.

Explanation

The optimal complexity is not a universal constant — it depends on the specific dataset's size, noise level, and the true underlying function's complexity, which is exactly why practitioners use validation curves / learning curves and cross-validation to find the right complexity for their specific problem rather than hard-coding a model size. The 3-panel picture is a mnemonic, not a formula: it just visually anchors 'wiggly line hitting every point = overfitting = high variance' vs. 'flat line ignoring curvature = underfitting = high bias' vs. 'smooth curve tracking the trend = the target.'

Q43Bias, Variance & Generalization

Regarding the relationship between VC-dimension and the bias-variance tradeoff, which statements are correct?

Select all statements you believe are true, then check your answer.

Explanation

VC-dimension is the formal/theoretical lens on the same phenomenon the wiggly-vs-straight-line pictures illustrate informally — it quantifies 'how many different labelings can this hypothesis class realize,' and a class that can realize more labelings (higher VC-dim) is, all else equal, more capable of fitting noise in any specific training sample (higher variance), while a class that can realize fewer labelings (lower VC-dim) is more constrained and thus more stable but potentially unable to capture real structure (higher bias). This connects the intuitive bias-variance story directly to the rigorous generalization bound covered earlier.

Q44Bias, Variance & Generalization

Which statements are true about the difference between 'irreducible error' (sometimes called Bayes error / noise) and the bias/variance components of prediction error?

Select all statements you believe are true, then check your answer.

Explanation

Although the deck's slides focus on bias and variance, the fuller statistical-learning picture (expected test error = bias² + variance + irreducible error) is exactly the kind of extension an interviewer will push you toward. Irreducible error is a property of the problem itself — e.g. two patients with identical measured features can still have different outcomes due to factors you never measured — and it sets a hard floor on achievable test error that no amount of modeling cleverness can beat. Bias and variance, by contrast, are properties of your chosen model and data, and are the two levers you actually control.

Q45Bias, Variance & Generalization

Suppose you plot training error and test error against increasing model complexity, and you notice the two curves are already diverging significantly even at low complexity, with test error much worse than training error throughout. What would you check?

Select all statements you believe are true, then check your answer.

Explanation

An unusually persistent train/test gap even at low model complexity is a strong signal to look beyond the bias-variance framing entirely, toward data or pipeline problems: distribution shift between train and test (the model learned a real pattern that just doesn't hold on the test distribution), evaluation bugs (shuffled labels, wrong preprocessing applied at test time), or a test set too small to be a reliable estimate (high variance in the error estimate itself, distinct from model variance). Jumping straight to 'the model underfits' without ruling these out is a common and costly mistake in real debugging.

Q46Bias, Variance & Generalization

Which statements are true regarding how dataset size interacts with the bias-variance tradeoff?

Select all statements you believe are true, then check your answer.

Explanation

This is the practitioner's mental model behind reading learning curves (error vs. training-set-size), a very common interview whiteboard exercise: variance shrinks as N grows because the sampled training set becomes a more faithful representative of the true distribution, but bias is essentially fixed by model capacity choice, so a genuinely underfit model plateaus at high error regardless of data size — the fix there is more capacity/better features, not more rows. An overfitting model, in contrast, usually DOES benefit from more data, since more diverse examples make it harder to memorize noise, even without touching model capacity.

Q91Bias, Variance & Generalization

An ensemble method (like a random forest) averages predictions from many individually high-variance, low-bias decision trees. Which statements correctly explain why this tends to reduce overall variance without proportionally increasing bias?

Select all statements you believe are true, then check your answer.

Explanation

This is the statistical mechanism behind why bagging/random forests work, and it's a very natural follow-up once someone understands bias/variance: averaging N independent, unbiased estimators reduces variance by roughly a factor of N (variance of the mean of iid variables), while leaving the average bias essentially unchanged — but that benefit specifically requires the individual errors to be at least partially uncorrelated. That's exactly why random forests inject randomness (bagging rows via bootstrap sampling, and randomly restricting the feature subset considered at each split) rather than just training the same tree algorithm N times on the same full dataset, which would produce nearly identical, highly correlated trees and little variance-reduction benefit.

Q92Bias, Variance & Generalization

Which statements correctly describe the role of a held-out validation set (distinct from both the training set and the final test set) in the context of bias-variance tradeoff management?

Select all statements you believe are true, then check your answer.

Explanation

This three-way split (train/validation/test) exists precisely because of the overfitting dynamics discussed earlier: if you pick hyperparameters by minimizing training error, you'll always be pulled toward maximal complexity (since training error is monotonically non-increasing in complexity); if you instead pick them by peeking at test-set performance, you've effectively 'fit' your hyperparameters to the test set too, inflating your reported number relative to true generalization. K-fold cross-validation is the standard fix for small datasets, where a single validation split would itself have high variance as an estimate.

Q93Bias, Variance & Generalization

Two data scientists debate: one says 'always use the most complex model available and rely on regularization to control overfitting,' the other says 'start simple and only add complexity when justified by validation performance.' Which statements are defensible, senior-level takes on this debate?

Select all statements you believe are true, then check your answer.

Explanation

This kind of open-ended debate question is specifically designed to see whether a candidate can reason about tradeoffs rather than reciting a rule. Both instincts have real merit and real failure modes: starting simple is faster to iterate on and easier to debug (and a simple baseline is invaluable for catching pipeline bugs before they're hidden by a complex model's flexibility), while 'complex + well-regularized' is exactly the recipe behind most modern deep learning successes when data and compute are abundant. A strong answer names the actual constraints (data size, latency SLA, interpretability requirement, compute budget) that would tip the decision either way for a specific project, rather than defending one side dogmatically.

Section 6 · 5 questions

Shattering & VC-Dimension

A formal way to talk about how expressive a hypothesis class is.

Q25Shattering & VC-Dimension

What is meant by 'shattering' a set of points, and which statements about it are correct?

Select all statements you believe are true, then check your answer.

Explanation

The definition only requires one favorable configuration of the N points to exist for which all 2^N labelings are achievable — it does not require every arrangement of N points to be shatterable. This subtlety is exactly what trips people up and is exactly what's tested by the '3 points, one bad arrangement vs. one good arrangement' example in the deck: straight lines can't shatter 3 collinear-ish points laid out one specific (unlucky) way, but they CAN shatter 3 points laid out a different way — and that single favorable arrangement is enough to say 'straight lines shatter 3 points.'

Q26Shattering & VC-Dimension

The deck shows that straight lines in a 2D plane can shatter 3 points (for a favorable arrangement) but cannot shatter any arrangement of 4 points. What is the correct conclusion, and why does the specific arrangement of 3 points matter?

Select all statements you believe are true, then check your answer.

Explanation

This is one of the most commonly mis-stated definitions in ML interviews. VC-dimension is defined via a best-case existential quantifier over arrangements ('does THERE EXIST an arrangement of N points that can be shattered'), combined with 'and NO arrangement of N+1 points can be shattered.' So a hypothesis class can fail to shatter some unlucky arrangements of N points and still have VC-dimension ≥ N, as long as one lucky arrangement works. For 2D lines: some arrangement of 3 points is fully shatterable (max), but no arrangement of 4 points is (2^4=16 labelings; a line can only realize a subset, notably not the classic 'XOR-like' diagonal labeling), so VC-dim = 3.

Q27Shattering & VC-Dimension

What is the VC-dimension of the hypothesis class of axis-aligned rectangles in a 2D plane, and why?

Select all statements you believe are true, then check your answer.

Explanation

This is a classic worked VC-dimension example precisely because it shows that having infinitely many candidate hypotheses (infinitely many rectangles by position/size) does NOT imply infinite VC-dimension — what matters is the maximum number of points whose label combinations can all be realized. With 4 points in a 'diamond' arrangement, you can independently include/exclude each extremal point via rectangle placement, hitting all 2^4=16 labelings; but with 5 points, by pigeonhole one point ends up strictly inside the convex hull of the others in every rectangle that contains them, so you can never label it negative while the surrounding points are positive. Hence VC-dim = 4.

Q28Shattering & VC-Dimension

According to the VC-dimension generalization bound shown in the slides (test error < training error + a complexity term that grows with D, the VC-dimension, and shrinks with N, the number of training samples), which statements are correct?

Select all statements you believe are true, then check your answer.

Explanation

This VC bound (from statistical learning theory / PAC learning) is the rigorous backbone behind the intuitive 'more complex models need more data' rule of thumb used throughout ML practice. It's a probabilistic upper bound, not an equality — it says test error is likely to be within training error plus a term that grows with model complexity (D) and shrinks with sample size (N). This is exactly the theoretical grounding for practical heuristics like 'don't use a 175B-parameter model on 500 rows' or 'if you scale up model capacity, scale up your training data too.'

Q89Shattering & VC-Dimension

Which of the following statements correctly rank these hypothesis classes by VC-dimension in 2D, from lowest to highest: (i) a fixed threshold on one feature (a single axis-aligned half-line), (ii) a general straight line, (iii) axis-aligned rectangles?

Select all statements you believe are true, then check your answer.

Explanation

This ranking exercise builds real intuition for how added geometric flexibility (more boundary parameters) tends to raise VC-dimension: 1 parameter (threshold) → VC-dim 1; 2 parameters (slope+intercept, i.e. a line) → VC-dim 3; 4 parameters (rectangle's 4 edges) → VC-dim 4. But this correlation is a rule of thumb, not a theorem — there exist pathological hypothesis classes with a single real-valued parameter yet infinite VC-dimension (a classic counterexample uses a parameter that encodes arbitrary bit patterns), which is exactly why 'count the parameters' is an unreliable shortcut for VC-dimension in general, even though it works for the well-behaved geometric families covered here.

Section 7 · 13 questions

Preparing Data for ML

Cleaning, missing values, and why feature scale quietly breaks distance-based and gradient-based methods.

Q47Preparing Data for ML

Which data quality problems are explicitly called out as common challenges when preparing real-world data for ML (noise, missing data, duplicates, inconsistency)?

Select all statements you believe are true, then check your answer.

Explanation

Real datasets routinely contain several of these simultaneously — the same student-records table in the slides shows noisy ages, missing GPAs, and inconsistent class-year labels all at once. In an interview, being able to name these categories specifically (rather than a vague 'the data was messy') signals that you've actually done hands-on data cleaning, and it also maps directly onto specific fixes: outlier/range checks for noise, imputation or deletion for missing values, and canonicalization/standard-vocabulary mapping for inconsistency.

Q48Preparing Data for ML

Which statements correctly describe strategies for handling missing values, per the slides (data deletion vs. imputation)?

Select all statements you believe are true, then check your answer.

Explanation

The deletion-vs-imputation choice is a real bias/variance-style tradeoff on the data side: deletion is simple and avoids introducing 'fake' values, but it shrinks your sample and can introduce selection bias if missingness isn't random (e.g. income is disproportionately unreported by high earners — dropping those rows skews the remaining data). Imputation (mean/median/mode, or a regression/kNN model trained to predict the missing field from other columns) preserves sample size but injects estimated rather than observed values, which understates true uncertainty if not handled carefully (e.g. via multiple imputation).

Q49Preparing Data for ML

Which statements correctly describe techniques for dealing with noisy data mentioned in the slides (binning, regression, outlier analysis)?

Select all statements you believe are true, then check your answer.

Explanation

All three are legitimate, complementary noise-reduction techniques: binning trades resolution for robustness (smoothing local fluctuations), regression-based smoothing assumes an underlying trend and denoises around it, and outlier analysis specifically targets extreme/implausible values (like Age=200 in the slide's example) using statistical rules (z-score, IQR) or domain-specific range checks. Noise is at least as common in numeric columns (measurement error, sensor glitches, data-entry typos) as in categorical ones — the 'Age = 200' and 'Age = -22' examples in the deck are both numeric.

Q50Preparing Data for ML

A hospital lab has two features: Hemoglobin (normal ≈ 13,500 mg/dL) and Creatinine (normal ≈ 0.7 mg/dL). Without scaling, why can raw absolute deviations mislead an analysis or a distance-based ML model?

Select all statements you believe are true, then check your answer.

Explanation

This hemoglobin/creatinine example is a perfect illustration of a very real production bug: raw-unit deviations conflate 'large number' with 'clinically/statistically important,' which is backwards here — a Creatinine change from 0.7 to 1.7 (a 2.4x relative increase, clinically significant for kidney function) looks numerically tiny next to a Hemoglobin swing of 1,000 mg/dL that might be within normal biological variance. Any algorithm relying on distances or magnitudes (k-NN, k-means, SVM with RBF kernel, gradient descent with unscaled features) will let the large-magnitude feature dominate purely due to units — this is a real bug pattern, not a cosmetic issue, and it's the standard justification for feature scaling before training.

Q51Preparing Data for ML

Which statements correctly describe z-score standardization, defined as f'_i = (f_i − μ_f) / σ_f ?

Select all statements you believe are true, then check your answer.

Explanation

This is the standard 'fit on train, transform on train/val/test using train's statistics' rule — recomputing μ and σ separately on the validation or test set is a classic data-leakage-adjacent bug (it lets test-set information subtly influence preprocessing, and it also means train and test features aren't on a genuinely comparable scale). Unlike min-max scaling, z-scores are unbounded — a value 5 standard deviations out stays a large number, it isn't clipped into [0,1]. And yes, a zero-variance (constant) feature breaks the formula outright — in practice you'd drop or special-case such a feature before standardizing.

Q52Preparing Data for ML

Which statements correctly describe min-max normalization, defined as f'_i = (f_i − min) / (max − min) ?

Select all statements you believe are true, then check your answer.

Explanation

Min-max scaling is bounded but fragile: because the transform is entirely defined by the two most extreme observed values, a single outlier (say, one data-entry error of Age=200) will crush every other legitimate value into a tiny sliver of the [0,1] range, destroying the resolution needed to distinguish 'normal' values from each other. Z-score standardization is comparatively more robust to this specific failure mode (though not immune), which is why the choice between them is itself a data-dependent decision — min-max when you truly know the bounds and want a fixed range (e.g. pixel intensities 0–255), z-score when the data may have unbounded outliers.

Q53Preparing Data for ML

Why does feature scale specifically matter for gradient-descent-based training (as opposed to, say, a decision tree), and which statements are correct?

Select all statements you believe are true, then check your answer.

Explanation

This is a very common 'why does X help' interview question. Gradient descent takes steps proportional to the gradient in each parameter's direction; if one feature's scale is 10,000x another's, the loss surface becomes a long, narrow valley, and a step size that's safe for the small-scale direction is far too small (slow convergence) or a step size safe for the large-scale direction overshoots wildly in the small one — scaling features to comparable ranges makes the surface closer to circular/well-conditioned, letting one learning rate work well for all parameters. Tree-based splits ('is Hemoglobin > 13,000?') are threshold comparisons that are unaffected by any monotonic transform of that one feature, which is exactly why trees/forests/GBMs are commonly used without feature scaling.

Q54Preparing Data for ML

Which of the following statements about handling inconsistent categorical data (e.g. 'Sr' vs. 'Senior' vs. 'Jr' meaning the same class years) are correct?

Select all statements you believe are true, then check your answer.

Explanation

This is an underrated, extremely common real-world bug: without canonicalization, one-hot encoding (or any categorical encoder) will silently create separate columns/embeddings for 'Sr' and 'Senior,' splitting what should be one category's statistical signal across two sparse, weaker features — the model doesn't 'know' they mean the same thing unless you tell it. This is exactly the kind of bug that degrades a model's performance without throwing any error, making it especially dangerous; a quick unique-value audit per categorical column before encoding is standard practice for exactly this reason.

Q55Preparing Data for ML

Which of the following statements correctly describe why 'garbage in, garbage out' is emphasized for ML, per the slide 'Good ML models can't be designed without good quality data'?

Select all statements you believe are true, then check your answer.

Explanation

This is a widely echoed piece of practitioner wisdom (often summarized as 'data-centric AI'): in many real projects, hours spent auditing labels and fixing systematic data issues beat hours spent tuning hyperparameters or trying bigger architectures, because a model can only ever be as good as the signal actually present in its training data. Some issues (a subtly biased labeling process, a sensor that silently started drifting) require domain expertise to even notice — and because production data pipelines keep ingesting new data, data quality is an ongoing monitoring concern, not a one-time cleanup task.

Q56Preparing Data for ML

You standardize (z-score) your training features and train a model. At serving time, a new data point arrives with a Hemoglobin value of 40,000 mg/dL (far beyond anything seen in training). Which statements are correct?

Select all statements you believe are true, then check your answer.

Explanation

This tests whether you understand that standardization is a linear rescaling, not a safety mechanism — it will happily compute a z-score of, say, +40 for a wildly anomalous input rather than flagging or clipping it. In a production system, that's precisely why serious ML pipelines add an explicit input-validation/anomaly-detection layer (range checks, schema validation, drift monitors) upstream or alongside the model, rather than assuming feature scaling alone will protect the model from garbage or out-of-distribution inputs.

Q94Preparing Data for ML

You're given a dataset where the 'income' feature is right-skewed (most values are modest, with a long tail of very high incomes). Which statements about preprocessing choices are correct?

Select all statements you believe are true, then check your answer.

Explanation

This is a common real-world nuance beyond the raw standardization/min-max formulas: z-score standardization only recenters and rescales (shifts mean to 0, std to 1) — it's a purely linear transform, so a right-skewed distribution stays exactly as skewed after standardization, just shifted/rescaled. A log transform, by contrast, is nonlinear and specifically compresses large values proportionally more than small ones, which is why it's the standard first move for monetary amounts, population counts, and other naturally multiplicative/right-skewed quantities before any further linear scaling is applied.

Q95Preparing Data for ML

Which statements correctly describe the risk of 'data leakage' during preprocessing, and how it connects to the train/validation/test split discipline?

Select all statements you believe are true, then check your answer.

Explanation

This is one of the most practically important and most commonly violated rules in real pipelines — computing μ/σ or min/max (or target encodings, or imputation values) on the combined dataset before splitting lets information about the validation/test set's distribution 'leak' into the training-time transform, giving an overly optimistic offline metric that won't hold up once the model meets genuinely unseen production data. This is exactly why libraries like scikit-learn enforce a fit-on-train / transform-on-everything-else API pattern (fit_transform on train, transform only on val/test) — the API design itself is a defense against this specific, easy-to-make mistake.

Q99Preparing Data for ML

A categorical feature 'city' has 500 distinct values, most appearing only once or twice in the training set. Which statements about preprocessing this feature are correct?

Select all statements you believe are true, then check your answer.

Explanation

High-cardinality categoricals are a very common real-world extension of the 'prepare your features properly' theme — this is a frequent live-coding/system-design topic at companies with lots of categorical business data (cities, product SKUs, merchant IDs). One-hot encoding 500 rare categories mostly produces near-constant, uninformative columns (curse-of-dimensionality-adjacent), which is why frequency encoding, target encoding (with proper cross-validation to avoid leakage), embedding layers (in deep models), or simply bucketing rare values into 'other' are the standard fixes.

Section 8 · 11 questions

Linear Regression

Simple and multiple regression, the least-squares objective, and reading the error surface.

Q57Linear Regression

For simple linear regression y = mx + c fit to (Work Experience, Gross Salary) data, which statements are correct?

Select all statements you believe are true, then check your answer.

Explanation

Choosing 'y = mx + c' as the model family is a textbook restriction bias — it excludes quadratic, exponential, or any other non-linear relationship outright, regardless of what the data might actually look like; the hypothesis space here is literally 'the set of all straight lines,' parameterized by (m, c). Once trained, predicting on brand-new x values (interpolation or, more riskily, extrapolation) is exactly the point of the model — 'only usable on training data' would make the model useless for its actual job.

Q58Linear Regression

Given several candidate straight lines fit to the same scatter of (experience, salary) points, how do we decide which one is 'best,' and which statements are correct?

Select all statements you believe are true, then check your answer.

Explanation

This is the setup for least-squares regression: for each candidate (m, c), you compute a per-point error — specifically the vertical distance y_actual − y_predicted (since x is treated as fixed/known and only y is being predicted, unlike, say, total-least-squares which also accounts for error in x) — sum (or average) it across all training points, and pick the (m, c) that minimizes that total. This turns 'best fit' from a subjective visual judgment into a well-defined optimization problem.

Q59Linear Regression

The least-squares objective function is J(m, c) = (1/2) Σ (y^(n) − s^(n))², where y^(n) = m·x^(n) + c is the model's prediction and s^(n) is the actual label. Which statements are correct?

Select all statements you believe are true, then check your answer.

Explanation

Squaring solves the cancellation problem (a +3 error and a −3 error would otherwise sum to 0, hiding real mistakes) and creates convexity in (m, c), which guarantees a unique global minimum reachable via calculus/gradient descent — this convexity is exactly why linear regression has a closed-form solution (the normal equations) unlike most nonlinear models. The 1/2 factor is purely a convenience — it exists so that the derivative of the squared term (which produces a factor of 2) cancels cleanly, leaving a tidier gradient expression; it does not shift the location of the minimum, since scaling an objective by any positive constant doesn't change where it's minimized.

Q60Linear Regression

Which statements correctly distinguish simple linear regression from multiple linear regression?

Select all statements you believe are true, then check your answer.

Explanation

The 'linear' in linear regression technically refers to linearity in the parameters, not necessarily in the raw features — this is a subtlety that trips people up and that's worth stating precisely in an interview: even polynomial regression (y = m1·x + m2·x² + c) is a 'linear model' in this sense, because it's linear in (m1, m2, c), even though the fitted curve is visibly nonlinear in x. Multiple linear regression is optimizable by either the closed-form normal equations OR iterative gradient descent — the gradient-descent update rule generalizes cleanly to any number of features (θ_i ← θ_i − η·∂J/∂θ_i for each parameter i).

Q61Linear Regression

For multiple linear regression, the per-parameter gradient update is θ_i ← θ_i − η·Σ_n (y^(n) − s^(n))·x_i^(n). Which statements are correct?

Select all statements you believe are true, then check your answer.

Explanation

η is the learning rate (a small positive hyperparameter controlling step size), not the sample count — conflating these is a common beginner mistake. The gradient naturally couples each parameter's update to its own feature's values, because ∂/∂θ_i of (θ·x) is exactly x_i. And yes: if predictions are systematically too high (positive error) on points where x_i is positive, the gradient step subtracts a positive quantity from θ_i, nudging predictions down for those points on the next iteration — this is literally the mechanism by which gradient descent 'corrects' its mistakes.

Q62Linear Regression

Which statements correctly describe what it means for regression parameters (m, c) to define a point on the 'error surface' J(m, c), and why minimizing this surface matters?

Select all statements you believe are true, then check your answer.

Explanation

Visualizing training as 'finding the bottom of a bowl in parameter space' is the standard mental model for gradient-based optimization, and for ordinary least-squares linear regression it's a mathematically exact bowl (a convex quadratic), which is why gradient descent (or the closed-form normal equations) reliably finds the global optimum. This convexity guarantee is special to linear regression (and a handful of other convex objectives like logistic regression, SVMs) — it does NOT extend to deep neural networks, whose loss surfaces are famously non-convex, riddled with local minima and saddle points, which is exactly why NN optimization needs momentum, adaptive learning rates, and careful initialization in ways plain linear regression never does.

Q63Linear Regression

A least-squares linear regression is fit on a dataset that includes one extreme outlier point far away from the rest. Which statements correctly describe the consequence?

Select all statements you believe are true, then check your answer.

Explanation

This is a very standard interview follow-up once someone explains squared error: because error is squared, a point 10x further from the line than a 'typical' point contributes roughly 100x more to the total loss, so the optimizer will happily sacrifice fit quality on many normal points to reduce that one point's huge squared contribution — this is precisely why outlier removal/robust scaling matters before OLS regression, and why robust alternatives (Huber loss, MAE-based regression, RANSAC) exist specifically to reduce this sensitivity.

Q64Linear Regression

Which statements correctly describe when you would choose multiple linear regression over simple linear regression in practice?

Select all statements you believe are true, then check your answer.

Explanation

The decision to add a feature should be driven by domain reasoning and residual diagnostics (are the model's errors correlated with some variable you left out?), not an arbitrary rule. Multicollinearity (correlated features) is a real practical concern for coefficient interpretability and stability — highly correlated features can make individual coefficient estimates unstable/hard to interpret — but it doesn't invalidate the model's predictive ability outright; regularization (ridge regression) is a common fix. And piling on more features without enough data is a classic bias-variance tradeoff: you may reduce bias (capturing real effects) while increasing variance (overfitting to noise in the extra features), especially with a small n relative to p.

Q65Linear Regression

Suppose a fitted linear regression line has residuals (y_actual − y_predicted) that show a clear curved (U-shaped) pattern when plotted against x, rather than looking like random scatter around zero. Which statements are correct?

Select all statements you believe are true, then check your answer.

Explanation

Residual analysis is one of the most practical regression diagnostics — patterns in residuals (curvature, funnel shapes indicating changing variance/heteroscedasticity, or trends over time) reveal exactly what a simple R² or MSE number hides: where and how the model is failing. A curved residual pattern is the classic sign of underfitting a nonlinear relationship with a linear model — the fix is either transforming features (adding polynomial terms, log transforms) or moving to a nonlinear model family. Random, structureless residuals around zero are what you want to see and are consistent with the linearity assumption holding.

Q66Linear Regression

Which statements correctly describe why we don't just find (m, c) by brute-force trying every possible value, and how gradient-based optimization connects to what's covered in the deck?

Select all statements you believe are true, then check your answer.

Explanation

Linear regression is one of the rare ML models with an exact closed-form solution (the normal equations, θ = (XᵀX)⁻¹Xᵀy), so gradient descent is a choice, not a necessity, for linear regression specifically — it becomes necessary in practice mainly when XᵀX is too large to invert efficiently (very high-dimensional features) or when using stochastic/mini-batch variants for huge datasets that don't fit in memory. This is worth knowing precisely because the deck teaches gradient descent as the general-purpose tool that does generalize to models (like deep neural nets) that have no closed-form solution at all.

Q96Linear Regression

Which statements correctly describe multicollinearity in multiple linear regression, where two or more features are highly correlated with each other?

Select all statements you believe are true, then check your answer.

Explanation

This distinction — prediction accuracy vs. coefficient interpretability — is one that trips up a lot of people who only think of regression as a prediction tool: even with severe multicollinearity, a linear model can still predict well on new data drawn from the same distribution (since the combined effect of the correlated features is well-estimated even if it's ambiguous how to split credit between them individually), but trying to interpret 'feature A's coefficient means X' becomes unreliable, because the optimizer could have assigned that same combined effect many different ways across the correlated features. Ridge regression's L2 penalty specifically discourages any one coefficient from growing arbitrarily large to compensate for a correlated partner, which stabilizes the individual estimates.

Section 9 · 13 questions

Gradient Descent Variants

Batch, stochastic and mini-batch descent — the tradeoffs that actually show up in training logs.

Q67Gradient Descent Variants

For the 1-parameter update rule w ← w − η·∂J(w)/∂w, which statements correctly explain why this moves w toward a minimum of J?

Select all statements you believe are true, then check your answer.

Explanation

This sign-flipping logic is the entire mechanism of gradient descent in one line: the gradient points in the direction of steepest increase, so subtracting a positive multiple of it moves you toward decrease — downhill — regardless of which side of the minimum you're currently on. The step size is explicitly proportional to the gradient's magnitude (scaled by η), which is exactly why gradient descent naturally takes big steps on steep parts of the loss surface and small steps near a flat minimum — the steps shrink automatically as you approach convergence, since the gradient itself shrinks toward zero there.

Q68Gradient Descent Variants

η (the learning rate) is a hyperparameter in gradient descent. Which statements correctly describe the consequences of choosing it poorly?

Select all statements you believe are true, then check your answer.

Explanation

This is a near-universal interview question because picking a good learning rate is one of the most consequential, failure-prone choices in practical deep learning. A too-large η can cause the loss to bounce around or blow up to NaN/infinity (divergence); a too-small η wastes enormous compute crawling toward the minimum. η is a hyperparameter set before/outside training (via grid search, learning-rate finders, or schedules) — it is not itself optimized by the gradient of J, which is precisely what distinguishes hyperparameters (set by the practitioner) from parameters (learned by the optimizer).

Q69Gradient Descent Variants

Which statements correctly describe Batch Gradient Descent (BGD), as defined by J(m, c) = (1/2)Σ_{n=1}^{N} (mx_n + c − s_n)² computed over ALL N training points each step?

Select all statements you believe are true, then check your answer.

Explanation

Batch GD is precise (the true full-batch gradient) but expensive and, per the deck's explicit critique, 'not data efficient when data is very similar' — if your dataset has lots of near-duplicate examples, computing the gradient over all of them every step is wasted work, since a much smaller subsample would already estimate the same gradient direction almost as well. This tension (expensive-but-exact vs. cheap-but-noisy) is exactly what motivates the SGD and mini-batch variants covered next.

Q70Gradient Descent Variants

Which statements correctly describe Stochastic Gradient Descent (SGD), where a parameter update is computed from just ONE randomly shuffled training example at a time?

Select all statements you believe are true, then check your answer.

Explanation

The apparent contradiction — 'each step is cheap, yet SGD is called not computationally efficient' — is resolved by hardware utilization: per-example updates can't exploit vectorized/parallel matrix operations as well as batched computation can, so despite doing 'less math' per step in a naive sense, SGD often runs slower in wall-clock time on modern hardware (GPUs) than well-batched alternatives, on top of the noisy, oscillating convergence path. Best practice re-shuffles the dataset every epoch (as the pseudocode in the deck shows: 'randomly shuffle samples' inside the loop) specifically to avoid the optimizer learning a spurious pattern tied to a fixed example order.

Q71Gradient Descent Variants

Which statements correctly describe Mini-batch Gradient Descent, which computes the gradient over a small batch of size u at a time?

Select all statements you believe are true, then check your answer.

Explanation

Mini-batch GD is what's actually used in the overwhelming majority of real deep learning training loops (batch sizes like 32, 64, 256, 512) precisely because it hits a practical sweet spot: enough averaging to tame SGD's noisiness, small enough to keep memory bounded and allow frequent updates, and sized to align well with GPU parallel throughput (hence the common power-of-2 convention). Setting u=N recovers batch GD (all data used per step), not SGD (single example per step) — u=1 recovers SGD instead, so the two extremes of the mini-batch spectrum ARE batch GD and SGD respectively.

Q72Gradient Descent Variants

Comparing all three variants (Batch, Stochastic, Mini-batch) on the axes of 'gradient accuracy per step' and 'compute cost per step,' which statements are correct?

Select all statements you believe are true, then check your answer.

Explanation

This accuracy-vs-cost tradeoff table is exactly the kind of summary an interviewer wants you to be able to draw from memory. On total FLOPs per epoch, batch GD and mini-batch GD (over enough batches to cover the dataset) actually do roughly the same total amount of arithmetic per epoch — the real difference is in update frequency (mini-batch updates parameters many times per epoch, batch GD only once) and in how well that arithmetic maps onto parallel hardware, not in total compute per epoch.

Q73Gradient Descent Variants

Which statements correctly describe why training data is randomly shuffled before each epoch in SGD/mini-batch GD, per the pseudocode 'randomly shuffle samples in the training dataset' at the start of each loop?

Select all statements you believe are true, then check your answer.

Explanation

This is a subtle-but-real practical bug source: if a dataset is sorted by class label (all class-0 examples first, then all class-1), an unshuffled mini-batch pass would feed the optimizer long runs of a single class, causing the parameters to swing toward fitting whichever class currently dominates the stream — a distorted, oscillating training signal totally different from what a properly mixed sample would produce. That's exactly why the deck's pseudocode places the shuffle step inside the 'loop until convergence,' i.e. every single epoch, not just once before training starts.

Q74Gradient Descent Variants

Regarding parameter initialization in gradient descent (the deck notes parameters 'can be all zeros' or 'can be random'), which statements are correct?

Select all statements you believe are true, then check your answer.

Explanation

This is a case where the 'right' initialization depends entirely on the model family — for linear/logistic regression (a convex bowl with one global minimum), zero-init works fine because there's nowhere else to converge to. But for a multi-neuron neural network layer, if every neuron starts with identical weights, they receive identical gradients throughout training and stay identical forever — effectively collapsing a layer with, say, 128 neurons down to the representational power of 1 — which is exactly why frameworks default to random initialization schemes (Xavier/Glorot, He initialization) for neural nets specifically. Initialization also affects convergence speed even when it doesn't change the final answer, since starting closer to the minimum means fewer steps are needed.

Q75Gradient Descent Variants

You observe that your training loss curve is oscillating wildly and occasionally spiking to very large values instead of steadily decreasing. Which of the following are plausible causes and fixes?

Select all statements you believe are true, then check your answer.

Explanation

This is a realistic training-log debugging scenario, and all three legitimate causes listed are the standard first things an ML engineer checks, roughly in this order: (1) learning rate — the single most common cause of divergence/oscillation, (2) if the optimizer is single-example SGD, some jitter is simply expected and mini-batching helps, (3) unscaled features can produce huge gradient magnitudes in certain directions, effectively acting like an implicitly huge learning rate for that parameter. An oscillating loss says nothing definitive about whether the model family is capable of fitting the data — that question can only be answered once optimization itself is stable.

Q76Gradient Descent Variants

For multiple linear regression with p features, the mini-batch gradient-descent parameter update for θ_i is θ_i ← θ_i − η Σ_{n∈batch} (y^(n) − s^(n))·x_i^(n). Which statements are correct?

Select all statements you believe are true, then check your answer.

Explanation

This ties the preprocessing section directly back to gradient descent mechanics: because the per-parameter gradient literally multiplies the error term by that feature's raw value (x_i), an unscaled large-magnitude feature mechanically produces larger gradient steps for its own coefficient — regardless of whether that feature is actually more predictive — which reintroduces the exact 'ill-conditioned, elongated loss surface' problem discussed earlier. This is precisely why the standard preprocessing pipeline order is: clean data → scale features → THEN run gradient descent.

Q97Gradient Descent Variants

Which statements correctly describe the relationship between the number of epochs, the learning rate schedule, and convergence in gradient-descent training, as commonly seen in practice?

Select all statements you believe are true, then check your answer.

Explanation

This links the mini-batch pseudocode's inner/outer loop structure to real training practice: one pass through all mini-batches = one epoch, and it's entirely standard to run many epochs. Learning-rate schedules (step decay, cosine annealing, warmup) exist precisely to get the best of both worlds — a large initial η for fast early progress across the broad part of the loss landscape, shrinking toward a small η for precise convergence near the minimum. And more epochs is explicitly NOT always better — past some point, additional epochs let the model keep fitting the training set's noise (exactly the overfitting story from the bias-variance section), which is why validation-loss monitoring (and early stopping) rather than a fixed epoch count is the standard practice.

Q98Gradient Descent Variants

Comparing plain (vanilla) gradient descent to momentum-based or adaptive-learning-rate variants (e.g. as used in modern deep learning), which statements are correct, based on the concepts introduced in this deck?

Select all statements you believe are true, then check your answer.

Explanation

This question deliberately looks forward from the deck's material to connect it with what a candidate would encounter next in a deep-learning-focused role: every modern optimizer (Momentum, RMSProp, Adam, AdamW) is still, at its core, computing gradients of a loss and using them to update parameters — what they add on top are smarter ways to set the effective step size per parameter and per iteration (e.g. accumulating a velocity term to smooth out noisy gradients, or adapting η per-parameter based on historical gradient magnitudes). Understanding plain batch/SGD/mini-batch deeply — including exactly why learning rate and feature scale matter — is precisely the prerequisite that makes those more advanced optimizers' design choices make sense rather than feel like arbitrary tricks.

Q100Gradient Descent Variants

Which statements correctly describe what happens to the least-squares linear regression objective's gradient as the model's parameters approach the true minimum?

Select all statements you believe are true, then check your answer.

Explanation

This is the calculus fact underlying the entire gradient-descent story: for a smooth convex function, the gradient vanishing is both necessary and sufficient to identify the (unique, global) minimum, which is exactly why 'set the gradient to zero and solve' is the closed-form approach for linear regression, and why gradient descent's steps naturally taper off as training converges rather than needing an explicit stopping rule based on step count alone — watching the gradient norm shrink toward zero is itself a legitimate convergence diagnostic used in real training loops.

Section 10 · 12 questions

Probability Foundations

Sample spaces, events, axioms, and the machinery every probabilistic ML model sits on top of.

Q77Probability Foundations

Which statements correctly describe why real-world processes are often modeled probabilistically even when, at some deeper physical level, they might be deterministic?

Select all statements you believe are true, then check your answer.

Explanation

This nuance is worth having ready for a 'why do we need probability in ML at all' question: a coin flip is a good example of 'practically random, but not fundamentally random' — with perfect knowledge of force, angle, air resistance, etc. it's Newtonian-deterministic, but no one actually tracks all that, so we treat the outcome as random. Deterministic behavior absolutely exists too (e.g. dropped objects reliably fall down given the same coarse assumptions) — the point isn't that everything is random, it's that whenever we choose to abstract away detail, probability becomes the right tool for describing what's left.

Q78Probability Foundations

For the experiment 'roll a fair six-sided die,' which statements about the sample space are correct?

Select all statements you believe are true, then check your answer.

Explanation

Sample space (S) is simply the exhaustive set of everything that could happen in one run of the experiment. Discrete-and-finite (a die), discrete-and-countably-infinite (number of customers arriving before 5pm — unbounded but countable), and continuous-and-uncountable (exact rainfall, a real number) are all valid sample spaces — this distinction matters because it determines whether you'll compute probabilities via simple counting/ratios (finite case) or via integrating a density function (continuous case).

Q79Probability Foundations

An event is defined as any subset of the sample space. For rolling a fair die, let event A = 'the outcome is less than 3'. Which statements are correct?

Select all statements you believe are true, then check your answer.

Explanation

The favorable-over-total formula is a common early trap: it's only mathematically justified under the assumption of equally likely outcomes (a fair die, a fair coin) — apply it to a loaded die or a biased coin and you'll get a wrong answer, since some outcomes are more probable than others and simple counting no longer reflects true probability. Single-outcome events (called 'elementary events') are perfectly valid — {4} is just as legitimate a subset of S as {1,2} is.

Q80Probability Foundations

Which statements correctly describe the frequentist definition of probability, P(A) = lim_{n→∞} n(A)/n, where n(A) is the number of times event A occurred in n trials?

Select all statements you believe are true, then check your answer.

Explanation

This is exactly why the frequentist framing is more general than the classical 'equally likely outcomes' formula: it works for ANY probability, fair or biased, discrete or (via careful limiting arguments) continuous, purely by observing how often something happens over many repetitions. In practice, of course, you can't literally run infinitely many trials — 10,000 coin flips give you a very good estimate of P(heads), with a margin of error that shrinks as n grows (this is the statistical foundation behind confidence intervals and, more broadly, why 'more data' generally gives better probability/parameter estimates in ML).

Q81Probability Foundations

Which of the following are the correctly stated axioms/properties of a probability measure P(·), per the Kolmogorov axioms referenced in the slides?

Select all statements you believe are true, then check your answer.

Explanation

These are the Kolmogorov axioms — the formal foundation that everything else in probability theory (including every ML model that outputs a probability, like softmax classifiers) is built on top of. Any valid probability, however it's computed (classical counting, frequentist estimation, a neural network's softmax output, a Bayesian posterior), must satisfy all of these simultaneously — non-negativity, total probability 1, and additivity over disjoint events. A negative probability is never valid under any interpretation; if a computed value comes out negative, that's a bug, not a valid edge case.

Q82Probability Foundations

Two events A and B are called 'disjoint' or 'mutually exclusive.' Which statements are correct?

Select all statements you believe are true, then check your answer.

Explanation

Disjointness is about whether the underlying sets overlap, not just whether the event descriptions sound different — 'even' ∩ 'greater than 3' = {4, 6}, a nonempty overlap, so those two events are NOT disjoint even though they're clearly distinct events. This distinction is exactly what determines whether you can use the simple additive rule P(A∪B)=P(A)+P(B) or need the general addition rule with the overlap subtracted — getting this wrong is one of the most common introductory-probability mistakes.

Q83Probability Foundations

The general addition rule states P(A ∪ B) = P(A) + P(B) − P(A ∩ B). Which statements correctly explain why the subtraction term is needed?

Select all statements you believe are true, then check your answer.

Explanation

This is the classic Venn-diagram intuition: the overlapping region gets counted once inside P(A) and once again inside P(B), so subtracting P(A∩B) removes exactly the duplicated count, leaving each part of the union counted exactly once. The disjoint additive rule is simply the special case where the overlap term is zero — they're not two separate formulas, one is a special case of the other. The idea generalizes to three or more sets via inclusion-exclusion (add singles, subtract pairwise overlaps, add back triple overlaps, and so on) — it's a foundational identity used, for instance, in exact computation of union probabilities in reliability/risk models.

Q84Probability Foundations

Which statements correctly describe the complement rule, P(A) = 1 − P(Aᶜ), and a scenario where it's the easier way to compute a probability?

Select all statements you believe are true, then check your answer.

Explanation

'At least one' / 'at least N' probability questions are the classic use-case for the complement trick — directly enumerating every combination that satisfies 'at least one head' across 3 flips means summing several cases (exactly 1, exactly 2, exactly 3 heads), while the complement ('all tails') is a single, trivially computable case, 1/8. This exact trick shows up in ML practice too — e.g. computing error rate as 1 − accuracy, or 'probability the ensemble's majority vote is wrong' via complementary counting over individual model errors.

Q85Probability Foundations

Which statements correctly describe a discrete vs. continuous random variable, using the examples given (weather condition vs. temperature)?

Select all statements you believe are true, then check your answer.

Explanation

The discrete-vs-continuous distinction determines the entire mathematical machinery you use downstream: discrete variables get probability mass functions (a table/list of P(X=x) values summing to 1), while continuous variables get probability density functions f(v), where P(X=v) for any single exact value v is actually 0 (you can only assign positive probability to an interval, via P(a

Q86Probability Foundations

For a continuous random variable X with probability density function f(v), which statements are correct?

Select all statements you believe are true, then check your answer.

Explanation

A subtle but important point often mis-stated: unlike a discrete probability (which IS bounded by 1), a probability density f(v) is not itself a probability and can exceed 1 at a point — e.g. a uniform distribution on the narrow interval [0, 0.1] has constant density f(v)=10 there, because it's the area under the curve (density × interval width) that must integrate to 1, not the height of the curve at any single point. This exact confusion (assuming densities are capped at 1) is a common source of bugs when people implement custom likelihood functions by hand.

Q87Probability Foundations

Why is the material on sample spaces, events, and probability axioms directly relevant to building ML models — e.g. Bayes decision theory or a softmax classifier's output?

Select all statements you believe are true, then check your answer.

Explanation

This is the 'why does this matter' payoff of the probability-foundations material: the requirement that softmax outputs sum to 1 and stay non-negative isn't an arbitrary design choice, it's the classifier being engineered to satisfy Kolmogorov's axioms so its outputs can be legitimately interpreted as P(class | input). Likewise, every reported test accuracy is really a frequentist point-estimate of a true underlying probability, complete with the same finite-sample uncertainty caveats — which is exactly why practitioners report confidence intervals or run multiple evaluation seeds rather than trusting a single accuracy number at face value.

Q101Probability Foundations

Which of the following are correct examples of applying the addition rule and complement rule together in a practical ML-evaluation context?

Select all statements you believe are true, then check your answer.

Explanation

This closes the probability section by tying the addition/complement rules directly back to metrics an ML engineer computes daily: false-positive and false-negative are mutually exclusive outcomes for a single prediction (a given example is either one or the other, never both), so their probabilities add directly, while 'at least one of several independent things happens' is the textbook complement-rule shortcut. The caveat about needing good underlying probability estimates is a genuinely important, often-skipped point: these identities are mathematically exact given true probabilities, but in practice P(false positive) etc. are themselves estimated from a finite validation set and inherit that estimate's own uncertainty.