What Does Machine Learning Code Look Like?

By upGrad

Updated on Sep 30, 2026 | 8 min read | 3.47K+ views

Share:

Key Highlights

  • Machine learning code is short. A working model can fit in about 15 lines of Python.
  • Almost every script follows five steps: load data, prepare and split it, train a model, make predictions, and evaluate.
  • Six libraries cover most ML code: pandas, NumPy, Matplotlib, scikit-learn, TensorFlow, and PyTorch.
  • The task changes the details. Classification predicts a label, regression predicts a number, and clustering finds groups without labels.
  • Deep learning replaces fit() with a training loop, which gives you more control but adds more code.
  • To start, run a simple example in Google Colab, then swap in a different model and see what changes.
  • In this article, you will learn what machine learning code looks like, the five steps most scripts follow, the Python libraries behind them, and how the code changes for classification, regression, clustering, and deep learning.

Ready to write machine learning code yourself? Check our AI and Machine Learning course in India and learn Python, scikit-learn, and deep learning through hands-on projects. Get mentor support, build a real portfolio, and take your first step toward an ML career.

The Basic Structure of Machine Learning Code

Machine learning code is regular Python with a library doing the heavy lifting. Most scripts follow the same order: load data, prepare it, split it, train a model, make predictions, and measure the results.

Here is a small but real example. It predicts whether a customer will leave a company, using scikit-learn.

import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score

data = pd.read_csv("customers.csv")
X = data.drop(columns="churned")
y = data["churned"]

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

model = LogisticRegression()
model.fit(X_train, y_train)

predictions = model.predict(X_test)
print(accuracy_score(y_test, predictions))

This assumes the columns are already numeric. Real data rarely is, which is why cleaning takes so much room in most projects.

A few things are worth noticing here.

  • X holds the input columns, also called features. y holds the answer the model should learn to predict, called the label.
  • The data is split so the model is tested on rows it has never seen. Here, 80% is used for training and 20% for testing. Without this split, the score would only show how well the model memorized its own data.
  • Training is one line: model.fit(). Predicting is another. The math happens inside the library.
  • This pattern is common in scikit-learn, where nearly every model uses the same fit and predict methods. Swap LogisticRegression for a random forest and the rest of the code barely changes.

Also read: How to Implement Machine Learning Steps: A Complete Guide

Free Courses

Explore courses related to AI
Fundamentals of Deep Learning and Neural Networks
Fundamentals of Deep Learning and Neural Networks
13.9K+ learners
28 hrs of learning
Artificial Intelligence in the Real World
Artificial Intelligence in the Real World
7.12K+ learners
7 hrs of learning
ChatGPT for Developers
ChatGPT for Developers
1.05K+ learners
2 hrs of learning

Core Steps in Machine Learning Code

Most ML scripts have five steps. I'll keep using the customer churn example, where the goal is to predict who will cancel.

Simple infographic showing how machine learning code works through five steps: getting data, preparing data, training the model, making predictions, and checking performance.

1. Import Libraries and Load Data

Imports come first. Then you load your file.

import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report

df = pd.read_csv("customers.csv")

Now look at the data before you use it.

print(df.shape)
print(df.head())
df.info()
print(df.isna().sum())

This takes a minute. It shows how many rows you have, what type each column is, and where values are missing.

A number stored as text is a common problem. It is easy to fix now and annoying to find later.

2. Prepare and Split the Data

This step is usually the longest. Models can't handle blanks or text, so you clean both up.

Missing values can be dropped or filled in. Text columns, like a plan type, get turned into numbers. Numeric columns often get scaled so a column measured in thousands doesn't drown out one measured in single digits.

df = df.dropna()

X = df.drop(columns="churned")
y = df["churned"]
X = pd.get_dummies(X, drop_first=True)

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y
)

scaler = StandardScaler()
X_train = scaler.fit_transform(X_train)
X_test = scaler.transform(X_test)

Notice the order. The split happens first. The scaler learns from the training rows only, then gets applied to the test rows.

If you scale before splitting, the test data leaks into training. That is called data leakage. It makes your final score look better than it should.

stratify=y keeps the share of churned customers the same in both sets. random_state=42 makes the split repeatable.

3. Choose and Train the Model

Look at what you are predicting. A category, like churn or no churn, needs a classifier. A number, like a price, needs a regressor.

Start with a simple model. Its score becomes your baseline.

model = LogisticRegression(max_iter=1000)
model.fit(X_train, y_train)

Training is the fit() line. The model looks at the training data and works out how much each column matters.

Settings like max_iter are hyperparameters. You choose them yourself. The model doesn't learn them.

Want to try a different model? Swap in RandomForestClassifier() and leave the rest alone. Every scikit-learn model uses the same fit() and predict() calls.

4. Make Predictions

Now you give the model the test rows. It has never seen them.

preds = model.predict(X_test)
probs = model.predict_proba(X_test)[:, 1]

predict() returns a label for each customer, either 0 or 1. predict_proba() returns the chance of churn. The [:, 1] picks the column for the churn class.

The probability is often more useful. A team can call the customers with the highest risk first.

5. Evaluate the Model

The last step is a comparison. You check the predictions against the real answers.

print(classification_report(y_test, preds))

Don't rely on accuracy alone. Say only 5 in 100 customers churn. A model that always predicts "no churn" is 95% accurate. It also never finds a single customer who leaves.

The report gives you two more numbers. Precision shows how many flagged customers really left. Recall shows how many of the leavers you caught.

Regression problems use different metrics, such as MAE and RMSE.

One split can also be lucky. cross_val_score() tests the model on several splits, so you get a fairer picture. If the score is low, look at the data first. Cleaner data and better columns usually beat a fancier model.

Also read: Top 6 Machine Learning Solutions

AI Courses to upskill

Explore Artificial Intelligence Courses for Career Progression

Certification6 Months
Executive Post Graduate Certificate8 Months

What Does Machine Learning Code Look Like in Python?

In Python, machine learning code is short. A working model can fit in about 15 lines.

Here is a full script you can copy and run. It uses a dataset that ships with scikit-learn, so you don't need to download anything. The model predicts whether a tumor is benign or malignant.

from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score

X, y = load_breast_cancer(return_X_y=True)

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y
)

model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))
model.fit(X_train, y_train)

preds = model.predict(X_test)
print(accuracy_score(y_test, preds))

Run it and you should see an accuracy in the high 90s. That's a solid result for so little code.

Here is what to notice.

  • There are no loops and no math. Python libraries do the heavy work in the background. You mostly call functions and pass data between them.
  • The data lives in arrays. X is a table of numbers, and y is a list of answers. Libraries like NumPy and pandas store them, and every model expects that format.
  • The Pipeline bundles the scaler and the model into one object. When you call fit(), the scaler learns from the training data only. When you call predict(), it applies the same scaling to new rows. This also protects you from data leakage without extra effort.
  • The method names are always the same. Nearly every scikit-learn model uses fit() and predict(). Once you learn them, you can use hundreds of models.
  • Most people don't write this in a plain .py file at first. They use a notebook, like Jupyter or Google Colab. A notebook lets you run code one block at a time and see the output right below it. That's handy when you are still exploring the data.
  • Real projects look a bit different. The same steps get split into functions and files. You might have one for loading data, one for training, and one for saving the model. The logic stays the same, though.

Deep learning code is longer. Libraries like PyTorch and TensorFlow need you to define the network layers and, in PyTorch's case, write the training loop yourself. I'll cover that in a later section.

Want to lead AI, not just learn about it? The Executive Programme in Generative AI & Agentic AI for Leaders from IIIT Bangalore is built for business leaders. No coding needed. Work through 20+ real enterprise cases, build a board-ready AI Leadership Execution Dossier, and earn certificates from IIIT-B and Microsoft.

Common Python Libraries Used in Machine Learning

You don't need many libraries to start. Six cover most of what you will see in ML code.

1. pandas

pandas loads and cleans tabular data. It is usually the first library in a script.

import pandas as pd

df = pd.read_csv("customers.csv")
df = df.dropna()

Its main object is the DataFrame, which works like a spreadsheet in code.

2. NumPy

NumPy handles arrays and fast number crunching. Most other ML libraries use it under the hood.

import numpy as np

arr = np.array([1, 2, 3, 4])
print(arr.mean())

You won't always import it directly. But you will often see its arrays.

3. Matplotlib

Matplotlib draws charts. Use it to explore data and plot results.

import matplotlib.pyplot as plt

plt.hist(df["age"])
plt.show()

Seaborn sits on top of it and makes common charts easier.

4. scikit-learn

scikit-learn is the main library for classic machine learning. It covers models, preprocessing, and evaluation.

from sklearn.ensemble import RandomForestClassifier

model = RandomForestClassifier()
model.fit(X_train, y_train)

It works well for tabular data. Most beginners should start here.

5. TensorFlow

TensorFlow is Google's deep learning library. Its Keras interface keeps model code short.

from tensorflow import keras

model = keras.Sequential([
    keras.layers.Dense(32, activation="relu"),
    keras.layers.Dense(1, activation="sigmoid"),
])

It is common in production systems and mobile deployments.

6. PyTorch

PyTorch is Meta's deep learning library. It is popular in research and in most modern AI work.

import torch.nn as nn

model = nn.Sequential(
    nn.Linear(10, 32),
    nn.ReLU(),
    nn.Linear(32, 1),
)

You write the training loop yourself, which gives you more control.

Also read: How to Learn Machine Learning – Step by Step

How Machine Learning Code Varies by Task

The five steps stay the same. What changes is the model, the target, and the metric.

1. Classification

Use it when the answer is a category. Spam or not spam. Churn or no churn.

from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import f1_score

model = RandomForestClassifier(random_state=42)
model.fit(X_train, y_train)

preds = model.predict(X_test)
print(f1_score(y_test, preds))
  • Target: a label
  • Common models: logistic regression, random forest
  • Metrics: precision, recall, F1

2. Regression

Use it when the answer is a number. A house price. Next month's sales.

from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error

model = LinearRegression()
model.fit(X_train, y_train)

preds = model.predict(X_test)
print(mean_absolute_error(y_test, preds))

The code is almost identical to classification. Only the model and the metric change.

  • Target: a continuous value
  • Common models: linear regression, gradient boosting
  • Metrics: MAE, RMSE, R²

3. Clustering

Use it when you have no labels. You want the model to find groups on its own, such as customer segments.

from sklearn.cluster import KMeans

model = KMeans(n_clusters=3, random_state=42)
model.fit(X)

print(model.labels_)

Notice what is missing. There is no y and no train/test split. The model only sees X.

You also pick the number of clusters yourself. Scale your data first, since KMeans relies on distance.

  • Target: none
  • Common models: KMeans, DBSCAN
  • Metrics: silhouette score, plus a manual check that the groups make sense

4. Deep Learning

Use it for images, text, audio, and very large datasets. This is where the code looks most different.

import torch
import torch.nn as nn

model = nn.Sequential(nn.Linear(10, 32), nn.ReLU(), nn.Linear(32, 1))
loss_fn = nn.BCEWithLogitsLoss()
optimizer = torch.optim.Adam(model.parameters(), lr=0.001)

for epoch in range(20):
    optimizer.zero_grad()
    outputs = model(X_train)
    loss = loss_fn(outputs, y_train)
    loss.backward()
    optimizer.step()

This assumes X_train and y_train are already PyTorch tensors.

Instead of one fit() call, you write a loop. Each pass does the same five things:

  1. Clear old gradients
  2. Run the data through the network
  3. Measure the error
  4. Calculate how to fix it
  5. Update the weights

Real projects also feed data in small batches using a DataLoader. Keras hides the loop behind fit(), which is why many beginners start there.

  • Target: labels, numbers, or both
  • Common tools: PyTorch, TensorFlow
  • Metrics: loss curves, plus task metrics like accuracy

Also read: 5 Breakthrough Applications of Machine Learning

Conclusion 

Machine learning code is shorter than most people expect. A working model can fit in about 15 lines of Python. Almost every script follows the same path. Load the data, prepare it, train a model, predict, and evaluate.

Libraries do the heavy lifting. Only the model, target, and metric change from task to task. Most of your effort goes into the data, not the algorithm. To get started, run the tumor example in Google Colab. Then swap in a different model and see what changes.

Have doubts about your career path? Book a free consultation with our experts today.

Frequently Asked Questions (FAQs)

1. Do I need to know advanced math to write machine learning code?

No. Libraries like scikit-learn handle the math for you. Basic statistics and some algebra help you understand the results. You can build working models before you learn the theory.

2. Is machine learning code hard for beginners?

The code itself is not hard. Most scripts use simple function calls. The harder part is understanding your data and choosing the right approach. Start with small datasets and simple models.

3. Which programming language is best for machine learning?

Python is the most popular choice. It has the biggest set of ML libraries and community support. R is common in statistics and research. Julia and C++ show up in performance-heavy work.

4. How long does it take to learn to write machine learning code?

If you already know basic Python, you can build your first model in a few days. Getting comfortable with real projects usually takes a few months of regular practice.

5. Can I run machine learning code without a powerful computer?

Yes. Classic models run fine on a normal laptop. For heavy deep learning, free tools like Google Colab and Kaggle Notebooks give you cloud GPUs at no cost.

6. How is machine learning code different from regular code?

Regular code follows rules you write by hand. Machine learning code lets the model learn rules from examples. You spend less time writing logic and more time working with data.

7. What is the difference between machine learning and deep learning code?

Deep learning is a branch of machine learning that uses neural networks. Its code is usually longer and often needs a GPU. Classic ML code is shorter and works well on smaller, tabular datasets.

8. Can I use machine learning without coding?

Yes. No-code tools like Google's Teachable Machine, KNIME, and DataRobot let you build models with a visual interface. They are good for quick tests. Code gives you more control and flexibility.

9. Can I use a machine learning model in a real app or website?

Yes. Once a model is trained, you can save it and load it inside an app. Many teams wrap it in a small web service, called an API, so other software can send data and get predictions back. Cloud platforms like AWS, Google Cloud, and Azure also offer ready-made hosting for this.

10. What is overfitting, and how do I spot it in my code?

Overfitting happens when a model memorizes the training data and fails on new data. You see it when the training score is high and the test score is much lower. Simpler models, more data, and regularization can help.

11. Where can I find datasets to practice on?

Start with the sample datasets built into scikit-learn. Kaggle, the UCI Machine Learning Repository, and Hugging Face Datasets offer thousands more. Pick one that matches a problem you care about.

upGrad

1008 articles published

We are an online education platform providing industry-relevant programs for professionals, designed and delivered in collaboration with world-class faculty and businesses. Merging the latest technolo...

Speak with AI & ML expert

+91

By submitting, I accept the T&C and
Privacy Policy

India’s #1 Tech University

Executive Program in Generative AI for Leaders

76%

seats filled

View Program

Top Resources

Recommended Programs

LJMU

Liverpool John Moores University

Master of Science in Machine Learning & AI

Double Credentials

Master's Degree

18 Months

IIITB
bestseller

IIIT Bangalore

Executive Diploma in Machine Learning and AI

360° Career Support

Executive Diploma

12 Months

IIITB

IIIT Bangalore

Executive Programme in Generative AI & Agentic AI for Leaders

India’s #1 Tech University

Dual Certification

5 Months