Overview
Generative and discriminative models represent two different answers to “what should a model actually learn from labeled data?” A generative model learns the full joint distribution of inputs and labels — effectively how each class produces its data — while a discriminative model learns only the boundary needed to tell classes apart, without modeling how the data itself was produced. That difference drives everything from data efficiency to whether the model can create new examples.
Comparison Diagram
Comparison Table
| Aspect | Generative Model | Discriminative Model |
|---|---|---|
| What it models | Full joint distribution P(x, y) — how features and labels co-occur | Conditional distribution P(y|x) or a direct decision function |
| Learning goal | Model the process that generates data for each class | Model the boundary that best separates classes |
| Data requirements | Can leverage unlabeled data since P(x) alone is still useful | Needs labeled data; largely ignores unlabeled examples |
| Handling missing or incomplete inputs | Marginalizes over missing features using the joint distribution | Requires imputation or breaks without a complete feature vector |
| New sample generation | Can sample and generate realistic new data points | Cannot generate data, only classify or score existing inputs |
| Typical algorithms | Naive Bayes, HMM, GANs, VAEs, Gaussian Mixture Models | Logistic regression, SVM, CRF, most feedforward neural classifiers |
| Classification accuracy at scale | Often lower given abundant labeled data | Usually higher given abundant labeled data since it optimizes the task directly |
| Computational cost | Higher — must model the entire data distribution | Lower — only needs to learn the boundary |
Key Differences
- Generative models learn P(x,y), the joint distribution, while discriminative models learn P(y|x) directly
- Only generative models can synthesize entirely new data samples
- Discriminative models typically reach higher classification accuracy when labeled data is abundant
- Generative models can exploit unlabeled data and gracefully handle missing features
- Discriminative models are usually more computationally efficient since they skip modeling the full data distribution
When to Use Each
Generative Model
- Synthetic Data Generation: GANs and VAEs are generative models built specifically to sample new, realistic data resembling the training set.
- Semi-Supervised Learning: When labels are scarce but raw data is plentiful, generative models can use P(x) from unlabeled examples to improve estimates.
- Anomaly Detection: Modeling the normal data distribution lets you flag inputs with low probability under that distribution as anomalies.
- Missing Feature Robustness: Because the full joint distribution is modeled, missing inputs can be marginalized out rather than requiring imputation.
Discriminative Model
- High-Accuracy Classification: With enough labeled data, discriminative models optimize the classification task directly and typically outperform generative counterparts.
- Structured Prediction Tasks: Conditional random fields model P(y|x) directly for sequence labeling problems like named entity recognition without modeling how text is generated.
- Resource-Constrained Training: Skipping the full data distribution means fewer assumptions and often faster training for pure prediction tasks.