8.14. Troubleshooting

Situation

Snapshot creation fails.

Step 1

Check if the bricks are thinly provisioned by following these steps:

Execute the mount command and check the device name mounted on the brick path. For example:

# mount
/dev/mapper/snap_lvgrp-snap_lgvol on /rhgs/brick1 type xfs (rw)
/dev/mapper/snap_lvgrp1-snap_lgvol1 on /rhgs/brick2 type xfs (rw)

Run the following command to check if the device has a LV pool name.
```
lvs device-name
```
For example:
```
#  lvs -o pool_lv /dev/mapper/snap_lvgrp-snap_lgvol
   Pool
   snap_thnpool
```
If the Pool field is empty, then the brick is not thinly provisioned.
Ensure that the brick is thinly provisioned, and retry the snapshot create command.

Step 2

Check if the bricks are down by following these steps:

Execute the following command to check the status of the volume:
```
# gluster volume status VOLNAME
```
If any bricks are down, then start the bricks by executing the following command:
```
# gluster volume start VOLNAME force
```
To verify if the bricks are up, execute the following command:
```
# gluster volume status VOLNAME
```
Retry the snapshot create command.

Step 3

Check if the node is down by following these steps:

Execute the following command to check the status of the nodes:
```
# gluster volume status VOLNAME
```
If a brick is not listed in the status, then execute the following command:
```
# gluster pool list
```
If the status of the node hosting the missing brick is Disconnected, then power-up the node.
Retry the snapshot create command.

Step 4

Check if rebalance is in progress by following these steps:

Execute the following command to check the rebalance status:
```
gluster volume rebalance VOLNAME status
```
If rebalance is in progress, wait for it to finish.
Retry the snapshot create command.

Situation

Snapshot delete fails.

Step 1

Check if the server quorum is met by following these steps:

Execute the following command to check the peer status:
```
# gluster pool list
```
If nodes are down, and the cluster is not in quorum, then power up the nodes.
To verify if the cluster is in quorum, execute the following command:
```
# gluster pool list
```
Retry the snapshot delete command.

Situation

Snapshot delete command fails on some node(s) during commit phase, leaving the system inconsistent.

Solution

Identify the node(s) where the delete command failed. This information is available in the delete command's error output. For example:

# gluster snapshot delete snapshot1
Deleting snap will erase all the information about the snap. Do you still want to continue? (y/n) y
snapshot delete: failed: Commit failed on 10.00.00.02. Please check log file for details.
Snapshot command failed

On the node where the delete command failed, bring down glusterd using the following command:
```
# service glusterd stop
```
Important
If glusterd crashes, there is no functionality impact to this crash as it occurs during the shutdown. For more information, see Section 24.3, “Resolving glusterd Crash”
Delete that particular snaps repository in /var/lib/glusterd/snaps/ from that node. For example:
```
# rm -rf /var/lib/glusterd/snaps/snapshot1
```
Start glusterd on that node using the following command:
```
# service glusterd start.
```
Repeat the 2nd, 3rd, and 4th steps on all the nodes where the commit failed as identified in the 1st step.
Retry deleting the snapshot. For example:
```
# gluster snapshot delete snapshot1
```

Situation

Snapshot restore fails.

Step 1

Check if the server quorum is met by following these steps:

Execute the following command to check the peer status:
```
# gluster pool list
```
If nodes are down, and the cluster is not in quorum, then power up the nodes.
To verify if the cluster is in quorum, execute the following command:
```
# gluster pool list
```
Retry the snapshot restore command.

Step 2

Check if the volume is in Stop state by following these steps:

Execute the following command to check the volume info:
```
# gluster volume info VOLNAME
```
If the volume is in Started state, then stop the volume using the following command:
```
gluster volume stop VOLNAME
```
Retry the snapshot restore command.

Situation

The brick process is hung.

Solution

Check if the LVM data / metadata utilization had reached 100% by following these steps:

Execute the mount command and check the device name mounted on the brick path. For example:

# mount
      /dev/mapper/snap_lvgrp-snap_lgvol on /rhgs/brick1 type xfs (rw)
      /dev/mapper/snap_lvgrp1-snap_lgvol1 on /rhgs/brick2 type xfs (rw)

Execute the following command to check if the data/metadatautilization has reached 100%:

lvs -v device-name

For example:

#  lvs -o data_percent,metadata_percent -v /dev/mapper/snap_lvgrp-snap_lgvol
     Using logical volume(s) on command line
   Data%  Meta%
     0.40

Note

Ensure that the data and metadata does not reach the maximum limit. Usage of monitoring tools like Nagios, will ensure you do not come across such situations. For more information about Nagios, see Chapter 18, Monitoring Red Hat Gluster Storage

Situation

Snapshot commands fail.

Step 1

Check if there is a mismatch in the operating versions by following these steps:

Open the following file and check for the operating version:
```
/var/lib/glusterd/glusterd.info
```
If the operating-version is lesser than 30000, then the snapshot commands are not supported in the version the cluster is operating on.
Upgrade all nodes in the cluster to Red Hat Gluster Storage 3.2 or higher.
Retry the snapshot command.

Situation

After rolling upgrade, snapshot feature does not work.

Solution

You must ensure to make the following changes on the cluster to enable snapshot:

Restart the volume using the following commands.

# gluster volume stop VOLNAME
# gluster volume start VOLNAME

Restart glusterd services on all nodes.
```
# service glusterd restart
```

Learn

Try, buy, & sell

Communities

About Red Hat

Making open source more inclusive

About Red Hat Documentation

Theme

Red Hat legal and privacy links

Red Hat legal and privacy links