mirror of
https://github.com/ibm-s390-linux/s390-tools.git
synced 2026-08-05 02:14:52 +00:00
6324f62da7
With "zpcictl --reset DDDD:BB:FF.F" now causing a fully Linux driven reset where the Linux kernel does an explicit device driver unbind, disable and re-enable, let's also expose a way to instead have firmware perform a device reset by issuing an SCLP with SCLP_ERRNOTIFY_RESET. When firmware is done resetting the device it will then issue an error notification with PCI Error Code 0x3a indicating successful reset, which will subsequently cause the new kernel based automatic recovery mechanism to perform recovery in coordination with the device driver. This allows resetting devices without unbinding them from their device driver and thus without losing related block devices or network interfaces. This may also be used to test the automatic recovery mechanism. Reviewed-by: Matthew Rosato <mjrosato@linux.ibm.com> Signed-off-by: Niklas Schnelle <schnelle@linux.ibm.com> Signed-off-by: Jan Höppner <hoeppner@linux.ibm.com>
114 lines
4.2 KiB
Plaintext
114 lines
4.2 KiB
Plaintext
.\" Copyright IBM Corp. 2022
|
|
.\" s390-tools is free software; you can redistribute it and/or modify
|
|
.\" it under the terms of the MIT license. See LICENSE for details.
|
|
.\"
|
|
.\" Macro for inserting an option description prologue.
|
|
.\" .OD <long> [<short>] [args]
|
|
.de OD
|
|
. ds args "
|
|
. if !'\\$3'' .as args \fI\\$3\fP
|
|
. if !'\\$4'' .as args \\$4
|
|
. if !'\\$5'' .as args \fI\\$5\fP
|
|
. if !'\\$6'' .as args \\$6
|
|
. if !'\\$7'' .as args \fI\\$7\fP
|
|
. PD 0
|
|
. if !'\\$2'' .IP "\fB\-\\$2\fP \\*[args]" 4
|
|
. if !'\\$1'' .IP "\fB\-\-\\$1\fP \\*[args]" 4
|
|
. PD
|
|
..
|
|
.
|
|
.TH zpcictl 8 "Mar 2022" s390-tools zpcictl
|
|
.
|
|
.SH NAME
|
|
zpcictl - Manage PCI devices on IBM Z
|
|
.
|
|
.
|
|
.SH SYNOPSIS
|
|
.B "zpcictl"
|
|
.I "OPTIONS"
|
|
.I "DEVICE"
|
|
.
|
|
.
|
|
.SH DESCRIPTION
|
|
Use
|
|
.B zpcictl
|
|
to manage PCI devices on the IBM Z platform. In particular,
|
|
use this command to report defective PCI devices to the Support Element (SE).
|
|
|
|
.B Note:
|
|
For NVMe devices additional data (such as S.M.A.R.T. data) is collected and sent
|
|
with any error handling action. For this extendend data collection, the
|
|
smartmontools must be installed.
|
|
.PP
|
|
.
|
|
.
|
|
.SH DEVICE
|
|
A PCI slot address (e.g. 0000:00:00.0) or the main device node of an NVMe
|
|
device (e.g.
|
|
.I /dev/nvme0
|
|
).
|
|
.
|
|
.
|
|
.SH OPTIONS
|
|
.SS Error Handling Options
|
|
.OD reset "" "DEVICE"
|
|
Reset and re-initialize the PCI device and report a device error to the Support
|
|
Element (SE). The reset consists of a controlled shutdown and a subsequent
|
|
re-enabling of the device from the shut off state. This process destroys and
|
|
then re-creates higher level interfaces such as network interfaces and block
|
|
devices. This reset is disruptive and often requires manual intervention on
|
|
multiple layers. In particular, network interfaces that are part of a bonded
|
|
interface must be re-added to the bond after the reset. Similarly, block
|
|
devices backed by an NVMe that are part of a software RAID must be re-synced by
|
|
re-adding to the RAID after resetting the NVMe.
|
|
|
|
Use this reset option only if the less disruptive automatic recovery mechanism
|
|
is not supported by your kernel or it failed to restore the device's
|
|
functionality. Unsuccessful automatic recovery can result in kernel messages
|
|
indicating required manual intervention. If the device is malfunctioning
|
|
without automatic recovery being triggered, consider using the \fB--reset-fw\fR
|
|
option to trigger a less disruptive automatic recovery through
|
|
a firmware-driven reset.
|
|
.PP
|
|
.
|
|
.OD reset-fw "" "DEVICE"
|
|
Reset the PCI device using a firmware-driven reset that also reports a device
|
|
error on the Support Element (SE). If supported by your kernel, automatic recovery
|
|
re-initializes the device after the firmware reports a successful device reset.
|
|
|
|
Use this option if the device is malfunctioning and automatic recovery is
|
|
supported by the kernel but was not triggered. This condition can occur if the
|
|
error is not detected by the low level PCI interfaces. A successful automatic
|
|
recovery after the firmware-driven reset, is less disruptive than the full
|
|
reset that is performed by the \fB--reset\fR option. Other than the full reset,
|
|
the automatic recovery does not completely shut down the device and re-create
|
|
it from the shut down state. Instead, it works with the device driver to
|
|
restore the device in place. Thus, higher level interfaces such as network
|
|
interfaces and block devices remain intact. In particular, with this type of
|
|
reset high availability mechanisms like a bonded network interface or
|
|
a software RAID can transparently re-integrate the recovered device. For
|
|
example, after a failure and recovery, a software RAID can resync a stroage
|
|
device or a network interface can be re-integrated in a bond. In contrast to
|
|
a complete shut down, the device driver remains active and informs higher
|
|
layers of both the occurence of an error state and the eventual recovery.
|
|
.PP
|
|
.
|
|
.OD deconfigure "" "DEVICE"
|
|
Deconfigure the PCI device and prepare for any repair action. This action
|
|
changes the status of the PCI device from configured to reserved.
|
|
.PP
|
|
.
|
|
.OD report-error "" "DEVICE"
|
|
Report any device error for the PCI device.
|
|
The device is marked as defective but no further action is taken.
|
|
.PP
|
|
.
|
|
.SS General Options
|
|
.OD help "h" ""
|
|
Print usage information, then exit.
|
|
.PP
|
|
.
|
|
.OD version "v" ""
|
|
Print version information, then exit.
|
|
.PP
|