Explore Courses
Liverpool Business SchoolLiverpool Business SchoolMBA by Liverpool Business School
  • 18 Months
Bestseller
Golden Gate UniversityGolden Gate UniversityMBA (Master of Business Administration)
  • 15 Months
Popular
O.P.Jindal Global UniversityO.P.Jindal Global UniversityMaster of Business Administration (MBA)
  • 12 Months
New
Birla Institute of Management Technology Birla Institute of Management Technology Post Graduate Diploma in Management (BIMTECH)
  • 24 Months
Liverpool John Moores UniversityLiverpool John Moores UniversityMS in Data Science
  • 18 Months
Popular
IIIT BangaloreIIIT BangalorePost Graduate Programme in Data Science & AI (Executive)
  • 12 Months
Bestseller
Golden Gate UniversityGolden Gate UniversityDBA in Emerging Technologies with concentration in Generative AI
  • 3 Years
upGradupGradData Science Bootcamp with AI
  • 6 Months
New
University of MarylandIIIT BangalorePost Graduate Certificate in Data Science & AI (Executive)
  • 8-8.5 Months
upGradupGradData Science Bootcamp with AI
  • 6 months
Popular
upGrad KnowledgeHutupGrad KnowledgeHutData Engineer Bootcamp
  • Self-Paced
upGradupGradCertificate Course in Business Analytics & Consulting in association with PwC India
  • 06 Months
OP Jindal Global UniversityOP Jindal Global UniversityMaster of Design in User Experience Design
  • 12 Months
Popular
WoolfWoolfMaster of Science in Computer Science
  • 18 Months
New
Jindal Global UniversityJindal Global UniversityMaster of Design in User Experience
  • 12 Months
New
Rushford, GenevaRushford Business SchoolDBA Doctorate in Technology (Computer Science)
  • 36 Months
IIIT BangaloreIIIT BangaloreCloud Computing and DevOps Program (Executive)
  • 8 Months
New
upGrad KnowledgeHutupGrad KnowledgeHutAWS Solutions Architect Certification
  • 32 Hours
upGradupGradFull Stack Software Development Bootcamp
  • 6 Months
Popular
upGradupGradUI/UX Bootcamp
  • 3 Months
upGradupGradCloud Computing Bootcamp
  • 7.5 Months
Golden Gate University Golden Gate University Doctor of Business Administration in Digital Leadership
  • 36 Months
New
Jindal Global UniversityJindal Global UniversityMaster of Design in User Experience
  • 12 Months
New
Golden Gate University Golden Gate University Doctor of Business Administration (DBA)
  • 36 Months
Bestseller
Ecole Supérieure de Gestion et Commerce International ParisEcole Supérieure de Gestion et Commerce International ParisDoctorate of Business Administration (DBA)
  • 36 Months
Rushford, GenevaRushford Business SchoolDoctorate of Business Administration (DBA)
  • 36 Months
KnowledgeHut upGradKnowledgeHut upGradSAFe® 6.0 Certified ScrumMaster (SSM) Training
  • Self-Paced
KnowledgeHut upGradKnowledgeHut upGradPMP® certification
  • Self-Paced
IIM KozhikodeIIM KozhikodeProfessional Certification in HR Management and Analytics
  • 6 Months
Bestseller
Duke CEDuke CEPost Graduate Certificate in Product Management
  • 4-8 Months
Bestseller
upGrad KnowledgeHutupGrad KnowledgeHutLeading SAFe® 6.0 Certification
  • 16 Hours
Popular
upGrad KnowledgeHutupGrad KnowledgeHutCertified ScrumMaster®(CSM) Training
  • 16 Hours
Bestseller
PwCupGrad CampusCertification Program in Financial Modelling & Analysis in association with PwC India
  • 4 Months
upGrad KnowledgeHutupGrad KnowledgeHutSAFe® 6.0 POPM Certification
  • 16 Hours
O.P.Jindal Global UniversityO.P.Jindal Global UniversityMaster of Science in Artificial Intelligence and Data Science
  • 12 Months
Bestseller
Liverpool John Moores University Liverpool John Moores University MS in Machine Learning & AI
  • 18 Months
Popular
Golden Gate UniversityGolden Gate UniversityDBA in Emerging Technologies with concentration in Generative AI
  • 3 Years
IIIT BangaloreIIIT BangaloreExecutive Post Graduate Programme in Machine Learning & AI
  • 13 Months
Bestseller
IIITBIIITBExecutive Program in Generative AI for Leaders
  • 4 Months
upGradupGradAdvanced Certificate Program in GenerativeAI
  • 4 Months
New
IIIT BangaloreIIIT BangalorePost Graduate Certificate in Machine Learning & Deep Learning (Executive)
  • 8 Months
Bestseller
Jindal Global UniversityJindal Global UniversityMaster of Design in User Experience
  • 12 Months
New
Liverpool Business SchoolLiverpool Business SchoolMBA with Marketing Concentration
  • 18 Months
Bestseller
Golden Gate UniversityGolden Gate UniversityMBA with Marketing Concentration
  • 15 Months
Popular
MICAMICAAdvanced Certificate in Digital Marketing and Communication
  • 6 Months
Bestseller
MICAMICAAdvanced Certificate in Brand Communication Management
  • 5 Months
Popular
upGradupGradDigital Marketing Accelerator Program
  • 05 Months
Jindal Global Law SchoolJindal Global Law SchoolLL.M. in Corporate & Financial Law
  • 12 Months
Bestseller
Jindal Global Law SchoolJindal Global Law SchoolLL.M. in AI and Emerging Technologies (Blended Learning Program)
  • 12 Months
Jindal Global Law SchoolJindal Global Law SchoolLL.M. in Intellectual Property & Technology Law
  • 12 Months
Jindal Global Law SchoolJindal Global Law SchoolLL.M. in Dispute Resolution
  • 12 Months
upGradupGradContract Law Certificate Program
  • Self paced
New
ESGCI, ParisESGCI, ParisDoctorate of Business Administration (DBA) from ESGCI, Paris
  • 36 Months
Golden Gate University Golden Gate University Doctor of Business Administration From Golden Gate University, San Francisco
  • 36 Months
Rushford Business SchoolRushford Business SchoolDoctor of Business Administration from Rushford Business School, Switzerland)
  • 36 Months
Edgewood CollegeEdgewood CollegeDoctorate of Business Administration from Edgewood College
  • 24 Months
Golden Gate UniversityGolden Gate UniversityDBA in Emerging Technologies with Concentration in Generative AI
  • 36 Months
Golden Gate University Golden Gate University DBA in Digital Leadership from Golden Gate University, San Francisco
  • 36 Months
Liverpool Business SchoolLiverpool Business SchoolMBA by Liverpool Business School
  • 18 Months
Bestseller
Golden Gate UniversityGolden Gate UniversityMBA (Master of Business Administration)
  • 15 Months
Popular
O.P.Jindal Global UniversityO.P.Jindal Global UniversityMaster of Business Administration (MBA)
  • 12 Months
New
Deakin Business School and Institute of Management Technology, GhaziabadDeakin Business School and IMT, GhaziabadMBA (Master of Business Administration)
  • 12 Months
Liverpool John Moores UniversityLiverpool John Moores UniversityMS in Data Science
  • 18 Months
Bestseller
O.P.Jindal Global UniversityO.P.Jindal Global UniversityMaster of Science in Artificial Intelligence and Data Science
  • 12 Months
Bestseller
IIIT BangaloreIIIT BangalorePost Graduate Programme in Data Science (Executive)
  • 12 Months
Bestseller
O.P.Jindal Global UniversityO.P.Jindal Global UniversityO.P.Jindal Global University
  • 12 Months
WoolfWoolfMaster of Science in Computer Science
  • 18 Months
New
Liverpool John Moores University Liverpool John Moores University MS in Machine Learning & AI
  • 18 Months
Popular
Golden Gate UniversityGolden Gate UniversityDBA in Emerging Technologies with concentration in Generative AI
  • 3 Years
Rushford, GenevaRushford Business SchoolDoctorate of Business Administration (AI/ML)
  • 36 Months
Ecole Supérieure de Gestion et Commerce International ParisEcole Supérieure de Gestion et Commerce International ParisDBA Specialisation in AI & ML
  • 36 Months
Golden Gate University Golden Gate University Doctor of Business Administration (DBA)
  • 36 Months
Bestseller
Ecole Supérieure de Gestion et Commerce International ParisEcole Supérieure de Gestion et Commerce International ParisDoctorate of Business Administration (DBA)
  • 36 Months
Rushford, GenevaRushford Business SchoolDoctorate of Business Administration (DBA)
  • 36 Months
Liverpool Business SchoolLiverpool Business SchoolMBA with Marketing Concentration
  • 18 Months
Bestseller
Golden Gate UniversityGolden Gate UniversityMBA with Marketing Concentration
  • 15 Months
Popular
Jindal Global Law SchoolJindal Global Law SchoolLL.M. in Corporate & Financial Law
  • 12 Months
Bestseller
Jindal Global Law SchoolJindal Global Law SchoolLL.M. in Intellectual Property & Technology Law
  • 12 Months
Jindal Global Law SchoolJindal Global Law SchoolLL.M. in Dispute Resolution
  • 12 Months
IIITBIIITBExecutive Program in Generative AI for Leaders
  • 4 Months
New
IIIT BangaloreIIIT BangaloreExecutive Post Graduate Programme in Machine Learning & AI
  • 13 Months
Bestseller
upGradupGradData Science Bootcamp with AI
  • 6 Months
New
upGradupGradAdvanced Certificate Program in GenerativeAI
  • 4 Months
New
KnowledgeHut upGradKnowledgeHut upGradSAFe® 6.0 Certified ScrumMaster (SSM) Training
  • Self-Paced
upGrad KnowledgeHutupGrad KnowledgeHutCertified ScrumMaster®(CSM) Training
  • 16 Hours
upGrad KnowledgeHutupGrad KnowledgeHutLeading SAFe® 6.0 Certification
  • 16 Hours
KnowledgeHut upGradKnowledgeHut upGradPMP® certification
  • Self-Paced
upGrad KnowledgeHutupGrad KnowledgeHutAWS Solutions Architect Certification
  • 32 Hours
upGrad KnowledgeHutupGrad KnowledgeHutAzure Administrator Certification (AZ-104)
  • 24 Hours
KnowledgeHut upGradKnowledgeHut upGradAWS Cloud Practioner Essentials Certification
  • 1 Week
KnowledgeHut upGradKnowledgeHut upGradAzure Data Engineering Training (DP-203)
  • 1 Week
MICAMICAAdvanced Certificate in Digital Marketing and Communication
  • 6 Months
Bestseller
MICAMICAAdvanced Certificate in Brand Communication Management
  • 5 Months
Popular
IIM KozhikodeIIM KozhikodeProfessional Certification in HR Management and Analytics
  • 6 Months
Bestseller
Duke CEDuke CEPost Graduate Certificate in Product Management
  • 4-8 Months
Bestseller
Loyola Institute of Business Administration (LIBA)Loyola Institute of Business Administration (LIBA)Executive PG Programme in Human Resource Management
  • 11 Months
Popular
Goa Institute of ManagementGoa Institute of ManagementExecutive PG Program in Healthcare Management
  • 11 Months
IMT GhaziabadIMT GhaziabadAdvanced General Management Program
  • 11 Months
Golden Gate UniversityGolden Gate UniversityProfessional Certificate in Global Business Management
  • 6-8 Months
upGradupGradContract Law Certificate Program
  • Self paced
New
IU, GermanyIU, GermanyMaster of Business Administration (90 ECTS)
  • 18 Months
Bestseller
IU, GermanyIU, GermanyMaster in International Management (120 ECTS)
  • 24 Months
Popular
IU, GermanyIU, GermanyB.Sc. Computer Science (180 ECTS)
  • 36 Months
Clark UniversityClark UniversityMaster of Business Administration
  • 23 Months
New
Golden Gate UniversityGolden Gate UniversityMaster of Business Administration
  • 20 Months
Clark University, USClark University, USMS in Project Management
  • 20 Months
New
Edgewood CollegeEdgewood CollegeMaster of Business Administration
  • 23 Months
The American Business SchoolThe American Business SchoolMBA with specialization
  • 23 Months
New
Aivancity ParisAivancity ParisMSc Artificial Intelligence Engineering
  • 24 Months
Aivancity ParisAivancity ParisMSc Data Engineering
  • 24 Months
The American Business SchoolThe American Business SchoolMBA with specialization
  • 23 Months
New
Aivancity ParisAivancity ParisMSc Artificial Intelligence Engineering
  • 24 Months
Aivancity ParisAivancity ParisMSc Data Engineering
  • 24 Months
upGradupGradData Science Bootcamp with AI
  • 6 Months
Popular
upGrad KnowledgeHutupGrad KnowledgeHutData Engineer Bootcamp
  • Self-Paced
upGradupGradFull Stack Software Development Bootcamp
  • 6 Months
Bestseller
KnowledgeHut upGradKnowledgeHut upGradBackend Development Bootcamp
  • Self-Paced
upGradupGradUI/UX Bootcamp
  • 3 Months
upGradupGradCloud Computing Bootcamp
  • 7.5 Months
PwCupGrad CampusCertification Program in Financial Modelling & Analysis in association with PwC India
  • 5 Months
upGrad KnowledgeHutupGrad KnowledgeHutSAFe® 6.0 POPM Certification
  • 16 Hours
upGradupGradDigital Marketing Accelerator Program
  • 05 Months
upGradupGradAdvanced Certificate Program in GenerativeAI
  • 4 Months
New
upGradupGradData Science Bootcamp with AI
  • 6 Months
Popular
upGradupGradFull Stack Software Development Bootcamp
  • 6 Months
Bestseller
upGradupGradUI/UX Bootcamp
  • 3 Months
PwCupGrad CampusCertification Program in Financial Modelling & Analysis in association with PwC India
  • 4 Months
upGradupGradCertificate Course in Business Analytics & Consulting in association with PwC India
  • 06 Months
upGradupGradDigital Marketing Accelerator Program
  • 05 Months

Apache Kafka Architecture: Comprehensive Guide For Beginners [2024]

Updated on 25 October, 2024

6.43K+ views
10 min read

Before we delve into the details of the Apache Kafka architecture, it is pertinent to shed some light on why Kafka makes headlines in the first place. To begin with, Apache Kafka mainly finds use in real-time streaming data architectures for providing real-time analytics. Durable, fast, scalable, and fault-tolerant, Kafka’s publish-subscribe messaging system has use cases for things like tracking IoT sensor data or tracking service calls.

Companies like LinkedIn, Netflix, Microsoft, Uber, Spotify, Goldman Sachs, Cisco, PayPal, and many others employ Apache Kafka for processing real-time streaming data. For example, LinkedIn, where Kafka originated, uses it to track operational metrics and activity data. Likewise, for Netflix, Apache Kafka is the de-facto standard for its messaging, eventing, and stream processing needs. 

Learn Online software development training from the World’s top Universities. Earn Executive PG Programs, Advanced Certificate Programs, or Masters Programs to fast-track your career.

The utility of Apache Kafka is better appreciated with an understanding of the Apache Kafka architecture and its underlying components. So, let’s explore the details of Kafka’s architecture.

Fundamental Kafka Architecture Concepts

The following concepts are basic to understanding the Apache Kafka architecture:

1. Topics

Kafka topics define the channels through which data is streamed. Thus, producers publish messages to the topics, and consumers read messages from the topics they subscribe. There is no limitation on the number of topics created within a Kafka cluster, and a unique name identifies each topic.

2. Brokers

Brokers are servers in a Kafka cluster that work as containers and hold multiple topics with distinct partitions. A unique integer ID identifies brokers in a Kafka cluster, and a connection with any one of these brokers means connecting with the entire cluster. 

3. Partitions

Kafka topics are divided into many parts known as partitions. Partitions are separated in order and allow multiple consumers to read data from a particular topic parallelly. The partitions of a topic are distributed across several servers in the Kafka cluster, and each server manages the data and requests for its lot of partitions. Messages reach the broker and a key, and the key determines the partition to which the particular message will go. Hence, messages with the same key go to the same partition. In case the key is unspecified, the partition is decided following a round-robin approach. 

4. Replicas 

In Kafka, replicas are like partition backups to ensure no data loss in case of a planned shutdown or failure. In other words, replicas are copies of partitions.

5. Partition Offsets

Since messages or records in Kafka are assigned to partitions, each record is provided with an offset to specify its position within the partition. Thus, the offset value associated with a record helps in its easy identification within the partition. A partition offset holds meaning within that particular partition only, and since records are added to partition ends, older records will have lower offset values.

6. Producers

Kafka producers publish messages to one or more topics and send data to the Kafka cluster. As soon as a producer publishes a message to a Kafka topic, the broker receives the message and adds it to a specific partition. Then, producers can choose the partition where they want to publish their message.

7. Consumers and Consumer Groups

Consumers read messages from the Kafka cluster. When a consumer is ready to receive the message, the data is pulled from the broker. Consumers belong to a consumer group, and each consumer within a particular group is responsible for reading a subset of the partitions of every topic it is subscribed.

8. Leader and Follower

Every Kafka partition has one server playing the role of leader. The leader performs all the read-and-write tasks for that particular partition. On the other hand, the job of the follower is to replicate the leader’s data. When a leader in a specific partition fails, one of the follower nodes assumes the role of the leader. A partition can have none or many followers.

9. Kafka Cluster

A Kafka cluster consists of one or more servers that are called brokers. A broker is a container that can hold multiple topics with different partitions. A unique integer ID is used to identify brokers in the Kafka cluster. The main goal of a Kafka cluster is to spread workloads evenly across replicas and partitions. Kafka clusters can scale without interruption by adding or removing brokers.

The following diagram is a simplified presentation of the interrelationships between the Apache Kafka architecture components discussed above.

Source

Apache Kafka Cluster Architecture

Here’s a detailed look at the main Kafka architectural components:

1. Kafka Brokers

Kafka clusters typically contain multiple nodes known as brokers. The brokers maintain the load balance. Each Kafka broker can handle hundreds and thousands of reads and writes every second. A broker serves as the leader for one particular partition. The leader has one or several followers, with the data on the leader replicated across the followers of that particular partition. 

Followers need to stay updated with the leader’s data. The leader, in turn, keeps track of the followers that are in sync with it. If a follower does not catch up with the leader or is no longer alive, it is removed from the in-sync replica list associated with the particular leader. A new leader is elected from among the followers upon the leader’s death, and the ZooKeeper supervises the election. Since the brokers are stateless, the ZooKeeper maintains its cluster state. The nodes in a cluster send heartbeat messages to the ZooKeeper to inform the latter that they are alive.  

2. Kafka Producers

Kafka producers directly send data to the brokers that play the role of a leader for a particular partition. The brokers or nodes of the Kafka clusters help the producers send direct messages. They do so by answering requests for metadata on which servers are alive and the live status of the partition leaders of a topic, enabling the producer to direct its requests accordingly. The producer decides which partition it wants to publish messages. Messages in Kafka are sent in batches, called record batches. Producers collect messages in memory and send them in batches either after a fixed period has elapsed or after a certain number of messages have accumulated.

3. Kafka Consumers

Kafka consumers issue requests to brokers denoting the partitions it wants to consume. The consumer specifies the partition offset in its request and receives a piece of log (starting from the offset position) from the broker. A log contains the records for a configurable period known as the retention period.

Consumers may also re-consume data as long as the log contains the data. Kafka consumers work on a pull-based approach which means that the brokers do not immediately push data onto the consumers. Instead, first, consumers send requests to brokers signalling that they are ready to consume data. Hence, the pull-based system ensures that the consumers are not overwhelmed with messages and can catch up if they fall behind. 

Following is a simplified Apache Kafka architecture diagram:

Source

Learn more about Apache Kafka.

Apache Kafka API Architecture

Apache Kafka has four key APIs – the Streams API, Connector API, Producer API, and Consumer API. Let’s see what role each has to play in enhancing the capabilities of Apache Kafka:

1. Streams API

The Streams API of Kafka allows an application to process data using a streams processing algorithm. Using the Streams API, applications can consume input streams from one or several topics, process them with stream operations, produce output streams, and eventually send them to one or more topics. Thus, the Streams API facilitates the transformation of input streams to output streams.

2. Connector API

The Connector API of Kafka is helpful for building, running, and managing reusable producers and consumers that connect Kafka topics to existing data systems or applications. For instance, a connector to a relational database could capture all updates and make sure the changes are available within a Kafka topic.

3. Producer API

The Producer API of Kafka allows applications to publish a stream of records to Kafka topics.

4. Consumer API

The Consumer API of Kafka Allows applications to subscribe to Kafka topics. It also enables applications to process record streams that are produced to those Kafka topics.

Role of ZooKeeper in Kafka

ZooKeeper is responsible for managing the configuration, coordination, and synchronisation of Kafka brokers within a Kafka cluster architecture –

  • Cluster Coordination

ZooKeeper maintains information about the Kafka cluster architecture, including the metadata about topics, partitions, brokers, and consumers. It serves as a centralized registry where brokers register themselves and consumers discover the current state of the Kafka cluster. ZooKeeper enables coordination and synchronization among the various components in the cluster.

  • Leader Election

Kafka relies on ZooKeeper for leader election within each partition. Each partition in Kafka has one broker designated as the leader, responsible for handling all read and write operations for that partition. If the leader fails, ZooKeeper is used to elect a new leader from the available replicas. By providing a consistent view of the leader for each partition, ZooKeeper ensures high availability and fault tolerance.

  • Broker Registration and Health Monitoring

When a Kafka broker starts up, it registers itself with ZooKeeper, providing details such as its hostname and port. ZooKeeper maintains a live list of available brokers, which is crucial for metadata management and load balancing. ZooKeeper also monitors the health of brokers by regularly checking their status. If a broker becomes unresponsive or fails, ZooKeeper updates the broker list accordingly.

How is the heartbeat managed for Kafka Brokers?

The heartbeat for Kafka brokers is managed by the following parameters –

  • heartbeat.interval.ms

The expected time between heartbeats to the consumer coordinator when using Kafka’s group management facilities. Heartbeats are used to ensure that the consumer’s session stays active and to facilitate rebalancing when new consumers join or leave the group. 

  • session.timeout.ms

The timeout is used to detect client failures when using Kafka’s group management facility. The client sends periodic heartbeats to indicate its liveness to the broker. If no heartbeats are received by the broker before the expiration of this session timeout, then the broker will remove this client from the group and initiate a rebalance

  • bootstrap.servers

bootstrap.servers is a configuration parameter that specifies a list of host and port pairs that are the addresses of the Kafka brokers in a “bootstrap” Kafka cluster that a Kafka client connects to initially to bootstrap itself. A host and port pair uses as the separator, for example, localhost:9092. The bootstrap.servers parameter is used for the initial connection to the Kafka cluster, but after that, Kafka will return advertised.listeners info which is a list of IP addresses that can be used to connect to the Kafka brokers

Way Forward

If you are interested to know more about Big Data, check out our Advanced Certificate Programme in Big Data from IIIT Bangalore.

Here’s an overview of the program with some key highlights:

  • Executive PGP from IIIT Bangalore with certifications in Data Science and Cloud Infrastructure
  • Online sessions and live lectures with 400+ hours of content
  • 7+ case studies and projects
  • 14+ programming languages and tools
  • 360-degree career support
  • Peer and industry networking

Sign up for more details about the course!

Frequently Asked Questions (FAQs)

1. What is Kafka used for?

Apache Kafka is mainly used for building real-time streaming data pipelines and applications adapting to those data streams. It allows both storage and analysis of real-time and historical data through a combination of messaging, storage, and stream processing.

2. Is Kafka a framework?

Apache Kafka is an open-source software that provides a framework for storing, reading, and analysing streaming data. Since it is open-source, Kafka is free to use with many developers and users contributing towards new features, updates, and support for new users.

3. Why do we need Kafka streams?

Kafka Streams is a client library for building microservices and streaming applications where the input data and output data are stored in the Apache Kafka cluster. On the one hand, it offers the benefits of Apache Kafka’s server-side cluster technology. On the other, it simplifies writing and deploying standard Scala and Java applications on the client side.

RELATED PROGRAMS