1. 故障排除概述
在NBU备份系统运行过程中,可能会遇到各种故障和错误。有效的故障排除是确保备份系统稳定运行的关键。本章节将介绍NBU备份系统的常见故障及其解决方案。更多学习教程www.fgedu.net.cn
# /usr/openv/netbackup/bin/admincmd/bperror -S
Status Code: 0
Message: the requested operation was successfully completed
Status Code: 1
Message: the requested operation was partially successful
Status Code: 2
Message: the requested operation was unsuccessful
Status Code: 3
Message: the system cannot find the path specified
Status Code: 4
Message: the system cannot open the file specified
Status Code: 5
Message: the specified file cannot be read
Status Code: 6
Message: the specified file cannot be written
Status Code: 7
Message: the specified file cannot be deleted
Status Code: 8
Message: the system cannot create the specified file
Status Code: 9
Message: the specified file already exists
2. 常见错误及解决方案
NBU备份系统常见的错误包括备份失败、恢复失败、服务异常等。以下是一些常见错误及其解决方案。学习交流加群风哥微信: itpux-com
2.1 备份失败错误
# /usr/openv/netbackup/bin/admincmd/bpdbjobs -failed -hours 24
Job ID Type State Status Client Policy Schedule
——- ———- ——– ————— ————– ————— —————
12345 Backup Done Failed client1 FULL_BACKUP Full
12346 Backup Done Failed client2 INCR_BACKUP Differential
# 查看失败作业的详细信息
# /usr/openv/netbackup/bin/admincmd/bpjobinfo -jobid 12345 -details
Job ID: 12345
Job Type: Backup
State: Done
Status: Failed
Client: client1
Policy: FULL_BACKUP
Schedule: Full
Start Time: 04/02/2026 20:00:00
End Time: 04/02/2026 20:05:00
Status Code: 2
Status Message: the requested operation was unsuccessful
# 查看作业详细日志
# /usr/openv/netbackup/bin/admincmd/bpjobinfo -jobid 12345 -log
19:59:59 – Info bpbkar (pid=12345) Backup started
20:00:00 – Info bpbkar (pid=12345) Backup of client1 started
20:00:01 – Info bpbkar (pid=12345) Waiting for storage unit selection
20:00:02 – Info bpbkar (pid=12345) Storage unit STU1 selected
20:00:03 – Info bpbkar (pid=12345) Error bpbrm (pid=12346) could not connect to client1: Connection refused
20:05:00 – Info bpbkar (pid=12345) Backup failed: status 2
2.2 恢复失败错误
# /usr/openv/netbackup/bin/admincmd/bpdbjobs -failed -hours 24 | grep Restore
Job ID Type State Status Client Policy Schedule
——- ———- ——– ————— ————– ————— —————
12347 Restore Done Failed client1 FULL_BACKUP Full
# 查看恢复失败作业的详细信息
# /usr/openv/netbackup/bin/admincmd/bpjobinfo -jobid 12347 -details
Job ID: 12347
Job Type: Restore
State: Done
Status: Failed
Client: client1
Policy: FULL_BACKUP
Schedule: Full
Start Time: 04/02/2026 10:00:00
End Time: 04/02/2026 10:05:00
Status Code: 28
Status Message: no entity was found
# 查看恢复作业详细日志
# /usr/openv/netbackup/bin/admincmd/bpjobinfo -jobid 12347 -log
09:59:59 – Info bprestore (pid=12347) Restore started
10:00:00 – Info bprestore (pid=12347) Restore of client1 started
10:00:01 – Info bprestore (pid=12347) Looking for backup images
10:00:02 – Info bprestore (pid=12347) Error: no entity was found
10:05:00 – Info bprestore (pid=12347) Restore failed: status 28
2.3 服务异常错误
# /usr/openv/netbackup/bin/bpps
NB Processes
———–
root 12345 1 0 10:00 ? 00:00:00 /usr/openv/netbackup/bin/bprd
root 12346 1 0 10:00 ? 00:00:00 /usr/openv/netbackup/bin/bpcd
root 12347 1 0 10:00 ? 00:00:00 /usr/openv/netbackup/bin/nbfsd
root 12348 1 0 10:00 ? 00:00:00 /usr/openv/netbackup/bin/vnetd
# 检查缺失的服务
# /usr/openv/netbackup/bin/bpps | grep bpjava
# 启动缺失的服务
# /usr/openv/netbackup/bin/bp.start_all
Starting NetBackup services:
Starting bprd… already running.
Starting bpcd… already running.
Starting nbfsd… already running.
Starting vnetd… already running.
Starting bpjava-msvc… started.
3. 日志分析
日志分析是故障排除的重要手段,NBU系统生成了大量的日志文件,记录了系统运行状态和错误信息。
3.1 日志文件位置
# ls -la /usr/openv/netbackup/logs/
total 48
drwxr-xr-x 12 root root 4096 Apr 2 10:00 .
drwxr-xr-x 20 root root 4096 Apr 2 09:00 ..
drwxr-xr-x 2 root root 4096 Apr 2 10:00 admin
-rw-r–r– 1 root root 0 Apr 2 10:00 bpbkar
-rw-r–r– 1 root root 0 Apr 2 10:00 bpbrm
-rw-r–r– 1 root root 0 Apr 2 10:00 bpcd
-rw-r–r– 1 root root 0 Apr 2 10:00 bpext
-rw-r–r– 1 root root 0 Apr 2 10:00 bprd
-rw-r–r– 1 root root 0 Apr 2 10:00 bpdbm
-rw-r–r– 1 root root 0 Apr 2 10:00 vnetd
# 查看客户端日志目录
# ls -la /usr/openv/netbackup/logs/client/
total 16
drwxr-xr-x 2 root root 4096 Apr 2 10:00 .
drwxr-xr-x 12 root root 4096 Apr 2 10:00 ..
-rw-r–r– 1 root root 100 Apr 2 10:00 log.12345
3.2 日志级别设置
# /usr/openv/netbackup/bin/vxlogcfg -a -p NB -o All -s DebugLevel=6
# 验证日志级别设置
# /usr/openv/netbackup/bin/vxlogcfg -l -p NB -o All
Log Level: 6
# 查看详细日志
# /usr/openv/netbackup/bin/vxlogview -p NB -o All -t 1h
Apr 2 10:00:00 client1 NB 137 bpbkar Info: Backup started
Apr 2 10:00:01 client1 NB 137 bpbkar Info: Backup of client1 started
Apr 2 10:00:02 client1 NB 137 bpbkar Info: Waiting for storage unit selection
Apr 2 10:00:03 client1 NB 137 bpbkar Error: bpbrm could not connect to client1: Connection refused
Apr 2 10:05:00 client1 NB 137 bpbkar Info: Backup failed: status 2
3.3 日志分析工具
# /usr/openv/netbackup/bin/admincmd/bperror -l -hours 24
04/02/2026 10:05:00 client1 ERR bpbkar (pid=12345) Backup failed: status 2 (the requested operation was unsuccessful)
04/02/2026 10:05:00 client1 ERR bpbrm (pid=12346) could not connect to client1: Connection refused
# 使用vxlogview查看详细日志
# /usr/openv/netbackup/bin/vxlogview -p NB -o bpbkar -t 24h
Apr 2 10:00:00 client1 NB 137 bpbkar Info: Backup started
Apr 2 10:00:01 client1 NB 137 bpbkar Info: Backup of client1 started
Apr 2 10:00:02 client1 NB 137 bpbkar Info: Waiting for storage unit selection
Apr 2 10:00:03 client1 NB 137 bpbkar Error: bpbrm could not connect to client1: Connection refused
Apr 2 10:05:00 client1 NB 137 bpbkar Info: Backup failed: status 2
4. 网络问题排查
网络问题是NBU备份失败的常见原因,包括网络连接中断、防火墙阻止、网络延迟等。
4.1 网络连接测试
# ping master_server
PING master_server (192.168.1.10) 56(84) bytes of data.
64 bytes from master_server (192.168.1.10): icmp_seq=1 ttl=64 time=1.2 ms
64 bytes from master_server (192.168.1.10): icmp_seq=2 ttl=64 time=1.1 ms
64 bytes from master_server (192.168.1.10): icmp_seq=3 ttl=64 time=1.3 ms
# 测试客户端与媒体服务器的连接
# ping media_server
PING media_server (192.168.1.20) 56(84) bytes of data.
64 bytes from media_server (192.168.1.20): icmp_seq=1 ttl=64 time=1.5 ms
64 bytes from media_server (192.168.1.20): icmp_seq=2 ttl=64 time=1.4 ms
64 bytes from media_server (192.168.1.20): icmp_seq=3 ttl=64 time=1.6 ms
4.2 端口测试
# telnet master_server 1556
Trying 192.168.1.10…
Connected to master_server.
Escape character is ‘^]’.
# 测试媒体服务器端口
# telnet media_server 13782
Trying 192.168.1.20…
Connected to media_server.
Escape character is ‘^]’.
# 查看NBU使用的端口
# /usr/openv/netbackup/bin/bpgetconfig | grep -i port
CLIENT_PORT_WINDOW = 1024-65535
SERVER_PORT_WINDOW = 1024-65535
4.3 防火墙设置
# systemctl status firewalld
● firewalld.service – firewalld – dynamic firewall daemon
Loaded: loaded (/usr/lib/systemd/system/firewalld.service; enabled; vendor preset: enabled)
Active: active (running) since Sat 2026-04-02 09:00:00 CST; 1h ago
# 查看防火墙规则
# firewall-cmd –list-all
public (active)
target: default
icmp-block-inversion: no
interfaces: eth0
sources:
services: ssh dhcpv6-client
ports: 1556/tcp 13720/tcp 13782/tcp
protocols:
masquerade: no
forward-ports:
source-ports:
icmp-blocks:
rich rules:
# 添加NBU所需端口
# firewall-cmd –add-port=1556/tcp –permanent
# firewall-cmd –add-port=13720/tcp –permanent
# firewall-cmd –add-port=13782/tcp –permanent
# firewall-cmd –reload
5. 存储问题排查
存储问题是NBU备份失败的另一个常见原因,包括存储空间不足、存储设备故障、存储网络问题等。
5.1 存储空间检查
# /usr/openv/netbackup/bin/admincmd/nbdevquery -liststs -U
Storage Server Name: storage1
Storage Type: PureDisk
Media Server Name: media1
State: UP
# 检查磁盘池状态
# /usr/openv/netbackup/bin/admincmd/nbdevquery -listdp -U
Disk Pool Name: DP1
Storage Server: storage1
Storage Type: PureDisk
State: UP
Capacity: 10000 GB
Free Space: 1000 GB
Used Space: 9000 GB
# 检查存储单元详细信息
# /usr/openv/netbackup/bin/admincmd/nbdevconfig -getconfig -storage_server storage1 -stype PureDisk
Storage Server: storage1
Storage Type: PureDisk
Media Server: media1
Connection String: storage1:9090
User Name: admin
Password: ******
Disk Pool: DP1
5.2 存储设备故障排查
# /usr/openv/netbackup/bin/admincmd/nbdevquery -listdv -U
Disk Volume Name: DV1
Disk Pool Name: DP1
Status: UP
Capacity: 10000 GB
Free Space: 1000 GB
Used Space: 9000 GB
# 检查存储设备详细信息
# /usr/openv/netbackup/bin/admincmd/nbdevconfig -getconfig -diskvolume DV1 -diskpool DP1
Disk Volume: DV1
Disk Pool: DP1
Status: UP
Path: /storage/disk1
Capacity: 10000 GB
Free Space: 1000 GB
5.3 存储网络问题排查
# ping storage1
PING storage1 (192.168.2.10) 56(84) bytes of data.
64 bytes from storage1 (192.168.2.10): icmp_seq=1 ttl=64 time=1.0 ms
64 bytes from storage1 (192.168.2.10): icmp_seq=2 ttl=64 time=0.9 ms
64 bytes from storage1 (192.168.2.10): icmp_seq=3 ttl=64 time=1.1 ms
# 测试存储网络带宽
# iperf3 -c storage1 -t 30
Connecting to host storage1, port 5201
[ 4] local 192.168.1.20 port 50000 connected to 192.168.2.10 port 5201
[ ID] Interval Transfer Bandwidth Retr Cwnd
[ 4] 0.00-30.00 sec 3.36 GBytes 960 Mbits/sec 0 1.40 MBytes
6. 客户端问题排查
客户端问题包括客户端服务未运行、客户端配置错误、客户端资源不足等。
6.1 客户端服务状态检查
# /usr/openv/netbackup/bin/bpps
NB Processes
———–
root 12345 1 0 10:00 ? 00:00:00 /usr/openv/netbackup/bin/bpcd
root 12346 1 0 10:00 ? 00:00:00 /usr/openv/netbackup/bin/vnetd
# 启动客户端服务
# /usr/openv/netbackup/bin/bp.start_all
Starting NetBackup services:
Starting bpcd… started.
Starting vnetd… started.
# 验证客户端服务状态
# /usr/openv/netbackup/bin/bpps
NB Processes
———–
root 12345 1 0 10:00 ? 00:00:00 /usr/openv/netbackup/bin/bpcd
root 12346 1 0 10:00 ? 00:00:00 /usr/openv/netbackup/bin/vnetd
6.2 客户端配置检查
# /usr/openv/netbackup/bin/bpgetconfig -M client1
CLIENT_CONNECT_TIMEOUT = 300
CLIENT_READ_TIMEOUT = 300
MAX_JOBS_PER_CLIENT = 4
# 检查客户端与主服务器的通信
# /usr/openv/netbackup/bin/bpclntcmd -pn
Expecting response from server master_server
client1.fgedu.net.cn client1 192.168.1.50:1556
6.3 客户端资源检查
# free -h
total used free shared buff/cache available
Mem: 16G 8G 4G 512M 4G 7G
Swap: 8G 0B 8G
# 检查客户端CPU使用情况
# top -b -n 1 | head -20
top – 10:00:00 up 10 days, 2:30, 2 users, load average: 0.50, 0.40, 0.30
Tasks: 200 total, 1 running, 199 sleeping, 0 stopped, 0 zombie
%Cpu(s): 10.0 us, 2.0 sy, 0.0 ni, 87.0 id, 1.0 wa, 0.0 hi, 0.0 si, 0.0 st
MiB Mem : 16384.0 total, 8192.0 free, 4096.0 used, 4096.0 buff/cache
MiB Swap: 8192.0 total, 8192.0 free, 0.0 used. 11264.0 avail Mem
# 检查客户端磁盘空间
# df -h
Filesystem Size Used Avail Use% Mounted on
/dev/sda1 50G 20G 30G 40% /
/dev/sdb1 500G 200G 300G 40% /data
7. 服务器问题排查
服务器问题包括主服务器服务异常、媒体服务器服务异常、数据库问题等。
7.1 主服务器服务状态检查
# /usr/openv/netbackup/bin/bpps
NB Processes
———–
root 12345 1 0 10:00 ? 00:00:00 /usr/openv/netbackup/bin/bprd
root 12346 1 0 10:00 ? 00:00:00 /usr/openv/netbackup/bin/bpcd
root 12347 1 0 10:00 ? 00:00:00 /usr/openv/netbackup/bin/nbfsd
root 12348 1 0 10:00 ? 00:00:00 /usr/openv/netbackup/bin/vnetd
root 12349 1 0 10:00 ? 00:00:00 /usr/openv/netbackup/bin/bpjava-msvc
root 12350 1 0 10:00 ? 00:00:00 /usr/openv/netbackup/bin/bpdbm
# 启动主服务器服务
# /usr/openv/netbackup/bin/bp.start_all
Starting NetBackup services:
Starting bprd… started.
Starting bpcd… started.
Starting nbfsd… started.
Starting vnetd… started.
Starting bpjava-msvc… started.
Starting bpdbm… started.
7.2 媒体服务器服务状态检查
# /usr/openv/netbackup/bin/bpps
NB Processes
———–
root 12345 1 0 10:00 ? 00:00:00 /usr/openv/netbackup/bin/bpcd
root 12346 1 0 10:00 ? 00:00:00 /usr/openv/netbackup/bin/vnetd
root 12347 1 0 10:00 ? 00:00:00 /usr/openv/netbackup/bin/nbrmms
# 启动媒体服务器服务
# /usr/openv/netbackup/bin/bp.start_all
Starting NetBackup services:
Starting bpcd… started.
Starting vnetd… started.
Starting nbrmms… started.
7.3 数据库问题排查
# /usr/openv/netbackup/bin/nbdb_admin -status
Database [NBDB] is in ADMIN mode
Database [BMRDB] is in ADMIN mode
# 验证数据库连接
# /usr/openv/netbackup/bin/nbdb_ping
Database [NBDB] is alive and well on server master_server.
Database [BMRDB] is alive and well on server master_server.
# 检查数据库备份状态
# /usr/openv/netbackup/bin/nbdb_backup -online -dbn NBDB -verbose
Backup of NBDB to /usr/openv/db/stage/NBDB_backup_1234567890.bak successful.
8. 故障排除最佳实践
遵循以下最佳实践,可以提高故障排除的效率和准确性。更多学习教程公众号风哥教程itpux_com
8.1 故障排除流程
- 收集信息:收集错误日志、作业状态、系统状态等信息
- 分析问题:根据收集的信息分析问题原因
- 制定解决方案:根据问题原因制定解决方案
- 实施解决方案:实施解决方案并验证效果
- 记录解决方案:记录解决方案和问题原因,以便后续参考
8.2 常见故障解决方案
# 1. 备份失败,错误代码2
# 原因:连接被拒绝
# 解决方案:检查客户端服务是否运行,检查防火墙设置
# 2. 备份失败,错误代码28
# 原因:找不到备份实体
# 解决方案:检查备份策略配置,检查客户端名称是否正确
# 3. 备份失败,错误代码13
# 原因:文件系统已满
# 解决方案:清理磁盘空间,扩展文件系统
# 4. 备份失败,错误代码59
# 原因:网络连接中断
# 解决方案:检查网络连接,检查网络设备状态
# 5. 恢复失败,错误代码28
# 原因:找不到备份映像
# 解决方案:检查备份映像是否存在,检查客户端名称是否正确
8.3 预防措施
- 定期监控备份作业状态,及时发现和解决问题
- 定期检查存储容量,确保有足够的存储空间
- 定期检查网络状态,确保网络连接稳定
- 定期检查系统资源,确保服务器和客户端有足够的资源
- 定期更新NBU软件,修复已知bug
- 建立完善的监控和告警机制,及时发现问题
8.4 故障排除工具
- bperror:查看错误信息
- bpjobinfo:查看作业详细信息
- vxlogview:查看详细日志
- nbdevquery:查看存储设备状态
- bpps:查看服务状态
- bpclntcmd:测试客户端与服务器的通信
# 1. 查看失败作业
# /usr/openv/netbackup/bin/admincmd/bpdbjobs -failed -hours 24
# 2. 查看作业详细信息
# /usr/openv/netbackup/bin/admincmd/bpjobinfo -jobid 12345 -details
# 3. 查看作业日志
# /usr/openv/netbackup/bin/admincmd/bpjobinfo -jobid 12345 -log
# 4. 查看错误信息
# /usr/openv/netbackup/bin/admincmd/bperror -l -hours 24
# 5. 测试网络连接
# ping master_server
# telnet master_server 1556
# 6. 检查服务状态
# /usr/openv/netbackup/bin/bpps
# 7. 检查存储状态
# /usr/openv/netbackup/bin/admincmd/nbdevquery -liststs -U
本文由风哥教程整理发布,仅用于学习测试使用,转载注明出处:http://www.fgedu.net.cn/10327.html
