Skip to main content

Manual backup and restore for Kafka

Platform backup was removed from Kafka on 15 September 2026, so protecting a Kafka cluster is now something you run yourself. This page sets up a dedicated mirror VM that continuously replicates your topics with MirrorMaker 2, snapshots them on a schedule, and restores a chosen snapshot onto a new Kafka service.

info

The Backup tab no longer appears on Kafka clusters, and the backup files that existed before that date were deleted. See Backup & Restore overview for what changed.

How it works​

The mirror VM is a plain VM you own, running a single-node Kafka plus MirrorMaker 2. It holds a live copy of your topics, which makes a consistent snapshot possible without ever stopping the source cluster.

PhaseWhat happens
MirrorMirrorMaker 2 continuously replicates every topic from the source cluster to the mirror VM
BackupA script stops Kafka on the mirror VM, archives the data directory, and restarts it
RestoreYou load a chosen archive onto the mirror VM, then replicate it to a new Kafka service with MirrorMaker 2

The source cluster is never stopped and never modified, in either direction.

Prerequisites​

ItemRequirement
OSUbuntu 20.04 or 22.04
Kafka binarykafka_2.13-<KAFKA_VERSION>.tgz, matching or compatible with the source cluster version
DiskTwo disks: one for the OS, one for data
NetworkThe mirror VM must reach the source cluster on port 9092

MirrorMaker 2 ships inside Kafka from version 2.4.0 onward, so there is nothing extra to install for it.

Placeholders used on this page​

PlaceholderMeaningExample
<KAFKA_VERSION>Kafka version to install3.8.0 or 4.1.2
<BACKUP_VM_IP>Private IP of the mirror VM10.0.1.50
<SOURCE_BOOTSTRAP>Bootstrap servers of the source cluster10.0.0.1:9092,10.0.0.2:9092
<SOURCE_ADMIN_USER>Admin username on the source clusteradmin
<SOURCE_ADMIN_PASS>Admin password on the source cluster—
<BACKUP_ADMIN_PASS>Admin password you set on the mirror VM—
<RESTORE_BOOTSTRAP>Bootstrap servers of the new Kafka service10.0.2.10:9092,10.0.2.11:9092
<RESTORE_ADMIN_USER>Admin username on the new Kafka serviceadmin
<RESTORE_ADMIN_PASS>Admin password on the new Kafka service—
<JAVA_HEAP_GB>JVM heap size in GB4
<DISK_DEVICE>Data disk device/dev/sdb on VMware, /dev/vdb on OpenStack
<CONFIG_PATH>Kafka config file, which moved between versionssee the table below

The config path and the controller quorum setting both changed across Kafka releases:

Kafka version<CONFIG_PATH>Controller quorum settingFormat command
Below 4.0/kafka/kafka/config/kraft/server.propertiescontroller.quorum.votersno extra flag
4.0 up to 4.1.2/kafka/kafka/config/kraft/server.propertiescontroller.quorum.bootstrap.serversadd --standalone
4.1.2 and above/kafka/kafka/config/server.propertiescontroller.quorum.bootstrap.serversadd --standalone

Part 1 — Install Kafka on the mirror VM​

Prepare the system​

# Update system packages
apt update -y

# Install Java (required by Kafka)
apt install -y default-jdk

# Verify Java installation
java -version

# Create Kafka system user and group
groupadd kafka
useradd -g kafka -s /bin/bash kafka

# Create base directories
mkdir -p /kafka

Prepare the data disk​

Replace <DISK_DEVICE> with the device name your platform uses — /dev/sdb on VMware, /dev/vdb on OpenStack.

# Partition the disk
parted <DISK_DEVICE> --script mklabel gpt
parted <DISK_DEVICE> --script mkpart primary 0% 100%
parted <DISK_DEVICE> --script set 1 lvm on

# Create LVM
pvcreate <DISK_DEVICE>1
vgcreate vg_data <DISK_DEVICE>1
lvcreate -l 100%FREE -n lv_data vg_data

# Format and mount
mkfs.xfs /dev/vg_data/lv_data
mkdir -p /data
mount /dev/vg_data/lv_data /data

# Persist mount across reboots
echo "/dev/mapper/vg_data-lv_data /data xfs users,noatime,auto,rw,nodev,exec,nosuid 0 0" >> /etc/fstab

# Set ownership
chown -R kafka:kafka /data

Install the Kafka binary​

Copy the Kafka package onto the VM, then extract it:

# Copy the Kafka binary package to the VM, then extract it
tar -xzf kafka_2.13-<KAFKA_VERSION>.tgz -C /kafka/
mv /kafka/kafka_2.13-<KAFKA_VERSION> /kafka/kafka

# Set ownership
chown -R kafka:kafka /kafka

Create the Kafka service​

cat > /lib/systemd/system/kafka.service << 'EOF'
[Unit]
Description=Apache Kafka
After=syslog.target network.target

[Service]
Type=simple
User=kafka
ExecStart=/kafka/kafka/bin/kafka-server-start.sh <CONFIG_PATH>
ExecStop=/kafka/kafka/bin/kafka-server-stop.sh
TimeoutStopSec=180
Restart=on-abnormal

[Install]
WantedBy=multi-user.target
EOF

systemctl daemon-reload

Substitute your version's <CONFIG_PATH> from the table above before running this.

Write the Kafka configuration​

Create the base config, a single node in KRaft mode:

# Create the config (single-node, KRaft mode)
# Replace <BACKUP_VM_IP> and <BACKUP_ADMIN_PASS> with actual values

cat > <CONFIG_PATH> << EOF
############################# KRaft Mode - Single Node #############################
process.roles=broker,controller
node.id=1

listener.security.protocol.map=CONTROLLER:SASL_PLAINTEXT,PLAINTEXT:PLAINTEXT,SASL_PLAINTEXT:SASL_PLAINTEXT
listeners=PLAINTEXT://localhost:9091,CONTROLLER://0.0.0.0:9094,SASL_PLAINTEXT://0.0.0.0:9092
advertised.listeners=PLAINTEXT://localhost:9091,SASL_PLAINTEXT://<BACKUP_VM_IP>:9092
inter.broker.listener.name=SASL_PLAINTEXT

############################# Controller #############################
EOF

Append the controller quorum setting that matches your version.

For Kafka below 4.0:

cat >> <CONFIG_PATH> << EOF
controller.quorum.voters=1@<BACKUP_VM_IP>:9094
EOF

For Kafka 4.0 and above:

cat >> <CONFIG_PATH> << EOF
controller.quorum.bootstrap.servers=<BACKUP_VM_IP>:9094
EOF

Then append the rest, which is the same on every version:

cat >> <CONFIG_PATH> << EOF
controller.listener.names=CONTROLLER

############################# SASL Authentication #############################
sasl.mechanism.controller.protocol=PLAIN
listener.name.controller.sasl.enabled.mechanisms=PLAIN
listener.name.controller.plain.sasl.jaas.config=org.apache.kafka.common.security.plain.PlainLoginModule required \
username="admin" \
password="<BACKUP_ADMIN_PASS>" \
user_admin="<BACKUP_ADMIN_PASS>";

listener.name.sasl_plaintext.plain.sasl.jaas.config=org.apache.kafka.common.security.plain.PlainLoginModule required \
username="admin" \
password="<BACKUP_ADMIN_PASS>" \
user_admin="<BACKUP_ADMIN_PASS>";

sasl.mechanism.inter.broker.protocol=PLAIN
sasl.enabled.mechanisms=PLAIN

############################# ACL #############################
authorizer.class.name=org.apache.kafka.metadata.authorizer.StandardAuthorizer
allow.everyone.if.no.acl.found=false
super.users=User:admin

############################# Storage & Log #############################
log.dirs=/data/kafka/combine-logs
num.partitions=1
offsets.topic.replication.factor=1
transaction.state.log.replication.factor=1
transaction.state.log.min.isr=1
default.replication.factor=1

log.retention.hours=168
log.retention.bytes=-1
log.segment.bytes=1073741824
log.retention.check.interval.ms=300000

############################# Network #############################
num.network.threads=3
num.io.threads=8
socket.send.buffer.bytes=102400
socket.receive.buffer.bytes=102400
socket.request.max.bytes=104857600
EOF

Set the JVM heap​

Use about 25% of the VM's RAM, at least 2 GB and at most 24 GB.

# Set heap size (adjust <JAVA_HEAP_GB> based on available RAM)
sed -i 's/export KAFKA_HEAP_OPTS="-Xmx[0-9]*G -Xms[0-9]*G"/export KAFKA_HEAP_OPTS="-Xmx<JAVA_HEAP_GB>G -Xms<JAVA_HEAP_GB>G"/' \
/kafka/kafka/bin/kafka-server-start.sh

Format storage and start Kafka​

Generate a cluster ID for the mirror VM. It is independent of the source cluster's ID.

CLUSTER_ID=$(/kafka/kafka/bin/kafka-storage.sh random-uuid)
echo "Cluster ID: $CLUSTER_ID"

Format the storage. For Kafka below 4.0:

/kafka/kafka/bin/kafka-storage.sh format \
-t $CLUSTER_ID \
-c <CONFIG_PATH>

For Kafka 4.0 and above:

/kafka/kafka/bin/kafka-storage.sh format \
-t $CLUSTER_ID \
-c <CONFIG_PATH> \
--standalone

Then start the service:

chown -R kafka:kafka /data

systemctl enable kafka
systemctl start kafka

# Verify Kafka is listening
ss -tlnp | grep 9092

Confirm Kafka works​

Create a credentials file for the CLI, then create and delete a throwaway topic:

# Create admin.properties for CLI access
cat > /kafka/kafka/config/kraft/admin.properties << EOF
security.protocol=SASL_PLAINTEXT
sasl.mechanism=PLAIN
sasl.jaas.config=org.apache.kafka.common.security.plain.PlainLoginModule required \
username="admin" \
password="<BACKUP_ADMIN_PASS>";
EOF

# Create a test topic to confirm Kafka is operational
/kafka/kafka/bin/kafka-topics.sh \
--bootstrap-server <BACKUP_VM_IP>:9092 \
--command-config /kafka/kafka/config/kraft/admin.properties \
--create \
--topic test-connectivity \
--partitions 1 \
--replication-factor 1

# List topics
/kafka/kafka/bin/kafka-topics.sh \
--bootstrap-server <BACKUP_VM_IP>:9092 \
--command-config /kafka/kafka/config/kraft/admin.properties \
--list

# Delete test topic
/kafka/kafka/bin/kafka-topics.sh \
--bootstrap-server <BACKUP_VM_IP>:9092 \
--command-config /kafka/kafka/config/kraft/admin.properties \
--delete \
--topic test-connectivity

Part 2 — Set up MirrorMaker 2​

MirrorMaker 2 replicates topics from the source cluster to the mirror VM. The configuration below uses IdentityReplicationPolicy, which keeps the original topic names instead of prefixing them — this is what makes the snapshot a faithful copy.

Write the MirrorMaker 2 configuration​

cat > /kafka/kafka/config/kraft/mm2.properties << EOF
# Cluster identifiers
clusters = source, target

# Source cluster (existing Kafka service)
source.bootstrap.servers = <SOURCE_BOOTSTRAP>
source.security.protocol = SASL_PLAINTEXT
source.sasl.mechanism = PLAIN
source.sasl.jaas.config = org.apache.kafka.common.security.plain.PlainLoginModule required \
username="<SOURCE_ADMIN_USER>" \
password="<SOURCE_ADMIN_PASS>";

# Target cluster (mirror VM)
target.bootstrap.servers = <BACKUP_VM_IP>:9092
target.security.protocol = SASL_PLAINTEXT
target.sasl.mechanism = PLAIN
target.sasl.jaas.config = org.apache.kafka.common.security.plain.PlainLoginModule required \
username="admin" \
password="<BACKUP_ADMIN_PASS>";

# Replication direction: source -> target
source->target.enabled = true
source->target.topics = .*
source->target.emit.checkpoints.enabled = false
source->target.sync.topic.acls.enabled = false

# Keep original topic names (no prefix added)
replication.policy.class = org.apache.kafka.connect.mirror.IdentityReplicationPolicy

# Internal topic replication factors (set to 1 for single-node mirror)
offset-syncs.topic.replication.factor = 1
heartbeats.topic.replication.factor = 1
checkpoints.topic.replication.factor = 1
mm2-offsets.topic.replication.factor = 1
config.storage.replication.factor = 1
status.storage.replication.factor = 1
offset.storage.replication.factor = 1
replication.policy.separator = .
replication.factor = 1
EOF

chown kafka:kafka /kafka/kafka/config/kraft/mm2.properties

source->target.topics = .* replicates everything. To mirror only certain topics, replace .* with a comma-separated list or a regex, such as orders,payments,user-events.

Create the MirrorMaker 2 service​

cat > /lib/systemd/system/kafka-mirror-maker.service << 'EOF'
[Unit]
Description=Kafka MirrorMaker 2 Service
After=network.target kafka.service

[Service]
Type=simple
User=kafka
ExecStart=/kafka/kafka/bin/connect-mirror-maker.sh /kafka/kafka/config/kraft/mm2.properties
StandardOutput=append:/kafka/kafka/logs/mm2.log
StandardError=append:/kafka/kafka/logs/mm2.log

[Install]
WantedBy=multi-user.target
EOF

systemctl daemon-reload

Rotate the MirrorMaker 2 log​

MirrorMaker 2 logs continuously, so rotate the file or it will fill the disk.

cat > /etc/logrotate.d/mm2 << 'EOF'
/kafka/kafka/logs/mm2.log {
maxsize 100M
copytruncate
rotate 2
delaycompress
compress
notifempty
missingok
su root root
}
EOF

# Run logrotate every 15 minutes
crontab -l 2>/dev/null | { cat; echo "15 * * * * /usr/sbin/logrotate /etc/logrotate.d/mm2"; } | crontab -

Start replication and verify it​

systemctl enable kafka-mirror-maker
systemctl start kafka-mirror-maker

# Check service status
systemctl status kafka-mirror-maker

# Monitor MM2 logs
tail -f /kafka/kafka/logs/mm2.log

After a few minutes, list the topics on the mirror VM. They should match the source cluster, with no prefix added:

/kafka/kafka/bin/kafka-topics.sh \
--bootstrap-server <BACKUP_VM_IP>:9092 \
--command-config /kafka/kafka/config/kraft/admin.properties \
--list

Part 3 — Take a backup​

A backup is a point-in-time archive of the mirror VM's data directory. The script stops MirrorMaker 2 and Kafka first, so the archive is internally consistent, then starts both again.

Create the backup script​

cat > /usr/local/bin/kafka-backup.sh << 'SCRIPT'
#!/bin/bash
set -e

KAFKA_DATA_DIR=/data/kafka/combine-logs
BACKUP_DIR=/backup/kafka
TIMESTAMP=$(date +%F_%H-%M-%S)
BACKUP_FILE=${BACKUP_DIR}/kafka-backup-${TIMESTAMP}.tar.gz
LOCKFILE=/tmp/kafka-backup.lock

# Prevent concurrent runs
if [ -f "$LOCKFILE" ]; then
echo "ERROR: Another backup is already running (lock file: $LOCKFILE). Exiting."
exit 1
fi
touch "$LOCKFILE"

cleanup() {
rm -f "$LOCKFILE"
}
trap cleanup EXIT

mkdir -p "$BACKUP_DIR"

echo "[$(date)] Starting Kafka backup..."

# Step 1: Stop MirrorMaker 2 to freeze incoming replication
echo "[$(date)] Stopping MirrorMaker 2..."
systemctl stop kafka-mirror-maker

# Step 2: Stop Kafka to ensure data consistency
echo "[$(date)] Stopping Kafka..."
systemctl stop kafka

# Step 3: Create compressed archive of the data directory
echo "[$(date)] Creating archive: $BACKUP_FILE"
tar -czf "$BACKUP_FILE" -C / data/kafka/combine-logs

echo "[$(date)] Backup size: $(du -sh $BACKUP_FILE | cut -f1)"

# Step 4: Restart Kafka
echo "[$(date)] Starting Kafka..."
systemctl start kafka
sleep 10

# Step 5: Restart MirrorMaker 2
echo "[$(date)] Starting MirrorMaker 2..."
systemctl start kafka-mirror-maker

echo "[$(date)] Backup completed successfully: $BACKUP_FILE"
SCRIPT

chmod +x /usr/local/bin/kafka-backup.sh

Run it, on a schedule or by hand​

Run a backup immediately:

/usr/local/bin/kafka-backup.sh

Schedule a daily backup at 02:00, and delete archives older than seven days:

crontab -l 2>/dev/null | { cat; echo "0 2 * * * /usr/local/bin/kafka-backup.sh >> /var/log/kafka-backup.log 2>&1"; } | crontab -

crontab -l 2>/dev/null | { cat; echo "30 2 * * * find /backup/kafka -name 'kafka-backup-*.tar.gz' -mtime +7 -delete"; } | crontab -

Change -mtime +7 to set a different retention period in days.

List what you have​

ls -lh /backup/kafka/

# Example output:
-rw-r--r-- 1 root root 2.3G 2026-04-10_02-00-01 kafka-backup-2026-04-10_02-00-01.tar.gz
-rw-r--r-- 1 root root 2.4G 2026-04-11_02-00-02 kafka-backup-2026-04-11_02-00-02.tar.gz
-rw-r--r-- 1 root root 2.4G 2026-04-12_02-00-01 kafka-backup-2026-04-12_02-00-01.tar.gz

Part 4 — Restore onto a new Kafka service​

Restoring has three phases: load the chosen archive onto the mirror VM, replicate it to a new Kafka service, then return the mirror VM to its normal job.

note

The mirror job from the source cluster is stopped for the duration. The source cluster itself is not touched.

Phase 1 — Load the archive onto the mirror VM​

Pick the archive you want and stop both services:

# List available backups
ls -lh /backup/kafka/

# Set the target backup filename
BACKUP_FILE=kafka-backup-<TIMESTAMP>.tar.gz

systemctl stop kafka-mirror-maker
systemctl stop kafka
danger

The next command permanently deletes the mirror VM's current data and replaces it with the archive. It cannot be undone. Confirm you selected the right timestamp before running it.

# Remove current data
rm -rf /data/kafka/combine-logs

# Extract the selected backup archive
tar -C /data -xzf /backup/kafka/$BACKUP_FILE --strip-components 1

# Fix ownership
chown -R kafka:kafka /data

systemctl start kafka
ss -tlnp | grep 9092

The mirror VM now serves exactly the data from the timestamp you chose.

Phase 2 — Replicate to the new service​

Create a new Kafka service in the portal, then note its bootstrap servers, admin username, and admin password. See Create a database.

Write a second MirrorMaker 2 configuration, this time from the mirror VM to the new service:

cat > /kafka/kafka/config/kraft/mm2-restore.properties << EOF
# Cluster identifiers
clusters = source, target

# Source cluster (mirror VM holding the point-in-time snapshot)
source.bootstrap.servers = <BACKUP_VM_IP>:9092
source.security.protocol = SASL_PLAINTEXT
source.sasl.mechanism = PLAIN
source.sasl.jaas.config = org.apache.kafka.common.security.plain.PlainLoginModule required \
username="admin" \
password="<BACKUP_ADMIN_PASS>";

# Target cluster (new Kafka service)
target.bootstrap.servers = <RESTORE_BOOTSTRAP>
target.security.protocol = SASL_PLAINTEXT
target.sasl.mechanism = PLAIN
target.sasl.jaas.config = org.apache.kafka.common.security.plain.PlainLoginModule required \
username="<RESTORE_ADMIN_USER>" \
password="<RESTORE_ADMIN_PASS>";

# Replication direction: source -> target
source->target.enabled = true
source->target.topics = .*
source->target.emit.checkpoints.enabled = false
source->target.sync.topic.acls.enabled = false

# Keep original topic names (no prefix added)
replication.policy.class = org.apache.kafka.connect.mirror.IdentityReplicationPolicy

# Replication factor = 1 (single-node mirror VM)
offset-syncs.topic.replication.factor = 1
heartbeats.topic.replication.factor = 1
checkpoints.topic.replication.factor = 1
mm2-offsets.topic.replication.factor = 1
config.storage.replication.factor = 1
status.storage.replication.factor = 1
offset.storage.replication.factor = 1
replication.policy.separator = .
replication.factor = 1
EOF

chown kafka:kafka /kafka/kafka/config/kraft/mm2-restore.properties

Create its service and start it:

cat > /lib/systemd/system/kafka-mirror-maker-restore.service << 'EOF'
[Unit]
Description=Kafka MirrorMaker 2 Restore Service
After=network.target kafka.service

[Service]
Type=simple
User=kafka
ExecStart=/kafka/kafka/bin/connect-mirror-maker.sh /kafka/kafka/config/kraft/mm2-restore.properties
StandardOutput=append:/kafka/kafka/logs/mm2-restore.log
StandardError=append:/kafka/kafka/logs/mm2-restore.log

[Install]
WantedBy=multi-user.target
EOF

systemctl daemon-reload

systemctl start kafka-mirror-maker-restore

# Monitor progress
tail -f /kafka/kafka/logs/mm2-restore.log

Check the topics landed on the new service:

cat > /kafka/kafka/config/kraft/restore-admin.properties << EOF
security.protocol=SASL_PLAINTEXT
sasl.mechanism=PLAIN
sasl.jaas.config=org.apache.kafka.common.security.plain.PlainLoginModule required \
username="<RESTORE_ADMIN_USER>" \
password="<RESTORE_ADMIN_PASS>";
EOF

/kafka/kafka/bin/kafka-topics.sh \
--bootstrap-server <RESTORE_BOOTSTRAP> \
--command-config /kafka/kafka/config/kraft/restore-admin.properties \
--list

The list should match what the archive contained. Once it does, stop the restore job:

systemctl stop kafka-mirror-maker-restore
systemctl disable kafka-mirror-maker-restore

Phase 3 — Return the mirror VM to normal​

Clear the snapshot, re-format the storage with a fresh cluster ID, and restart mirroring:

systemctl stop kafka

rm -rf /data/kafka/combine-logs

CLUSTER_ID=$(/kafka/kafka/bin/kafka-storage.sh random-uuid)
echo "Cluster ID: $CLUSTER_ID"

/kafka/kafka/bin/kafka-storage.sh format \
-t $CLUSTER_ID \
-c <CONFIG_PATH>

/kafka/kafka/bin/kafka-storage.sh format \
-t $CLUSTER_ID \
-c <CONFIG_PATH> \
--standalone

Format with the command for your version, exactly as in Part 1, then:

chown -R kafka:kafka /data

systemctl start kafka
sleep 10
systemctl start kafka-mirror-maker

MirrorMaker 2 starts re-syncing from the source cluster and the scheduled backups resume on their own.

Troubleshooting​

MirrorMaker 2 is not replicating​

Check the log, then confirm the mirror VM can actually reach the source cluster:

# Check MM2 logs for errors
tail -100 /kafka/kafka/logs/mm2.log | grep -i error

# Create source-admin.properties
cat > /kafka/kafka/config/kraft/source-admin.properties << EOF
security.protocol=SASL_PLAINTEXT
sasl.mechanism=PLAIN
sasl.jaas.config=org.apache.kafka.common.security.plain.PlainLoginModule required \
username="<SOURCE_ADMIN_USER>" \
password="<SOURCE_ADMIN_PASS>";
EOF

# Verify connectivity to the source cluster
/kafka/kafka/bin/kafka-topics.sh \
--bootstrap-server <SOURCE_BOOTSTRAP> \
--command-config /kafka/kafka/config/kraft/source-admin.properties \
--list

Kafka will not start after a restore​

The usual cause is a nested directory from the archive. /data/kafka/combine-logs/ should contain partition directories directly, not another combine-logs folder.

ls -la /data/kafka/combine-logs/
# Should show partition directories, not a nested combine-logs folder

# Check Kafka logs
journalctl -u kafka -n 50

The backup script stopped partway​

Kafka and MirrorMaker 2 may still be stopped, and the lock file may still be there. Start them and clear it:

systemctl start kafka
sleep 10
systemctl start kafka-mirror-maker

# Remove the stale lock file if present
rm -f /tmp/kafka-backup.lock

Next steps​