Posts

Tips for being a better system administrator

Here are a few tips that I have discovered or implemented during my career, that can help anyone get better results: Follow the 3-2-1 backup strategy  or better 3 copies of the data 3 different media 1 copy offsite Test your backups regularly. Ideally, automate your recovery tests. Use Whatever-as-code as much as you can. IaC is the first that comes to my mind, but using code to define objects and their properties have several advantages: Auto-documentation of changes Ability to easily rollback (using source code management like Git) Makes it easier for standardization an compliance Ansible, Chef, Puppet, Salt, Terraform, Pulumi are good examples Maintain a changelog of all the major changes in your infrastructure (that isn't already "documented" because you're using IaC). Keep track of your hardware inventory and plan your renewals and warranty extension purchases. Build a spreadsheet with all the components that you manage and make sure that everyone on the sysadmi...

CentOS6 - End of life - How to install packages?

 CentOS 6 is EOL (End of life) since November 2020.  How to install packages now that the default CentOS repos have been emptied? The script in this Github repo is very helpful for that:  https://github.com/EngineeredVirus/CentOS6-Repository-Fix . Have fun!

NRPE troubleshooting

 Hi,  To make nrpe daemon log plugins output, add this to the command definition: 2>&1   Example:    # 'check_local_mailq' command definition define command{         command_name    check_local_mailq         command_line    /usr/bin/sudo /usr/lib64/nagios/plugins/check_mailq -w 1 -c 2 > /tmp/nagiosdebug 2>&1         } Of course, make sure the user running nrpe or the plugin can write to the file ( /tmp/nagiosdebug in this case). Then: # tail -f /tmp/nagiosdebug sudo: sorry, you must have a tty to run sudo I just added this in my sudo configuration # Required for check_mailq Defaults:nagios !requiretty   And now it works! Important: once you're done debugging, remove the redirection, so that Nagios can get access to the output. Otherwise, you'll see "null" in the output instead of something like: OK: mailq i...

Symfony - clear cache

I was a bit surprised when I learnt that the only way to clear the cache on a Symfony 4 system is to run a local command on the system, as I don't want anyone to log into a production server (deployment is automated). And when you run the command, you must run it as the user that your web server is running under.  For distros of the Red Hat family, this user is apache by default.  For the Debian family, it is www-data. Here is how I did it.  Let's say that the user that is used for automated deployment is 'user123', here's what I did: Create a file in the /etc/sudoers.d directory (example: /etc/sudoers.d/user123_bin_console) Put this line in the file: user123 ALL=(apache) NOPASSWD: /path/to/bin/console Once this file is in place, you can log into the system and execute the sudo -l command to show what commands this user is allowed to execute via sudo and make sure the command you specified in the newly created file is present To clear the cache, execute this command...

Networker automated recovery testing using the REST API - first script

I have worked a lot on my automated recovery script recently and finally got to a version that is a really good proof of concept and that I'm not scared to publish.  Plus, I have created a Gitub repo for documentation, issue management and, of course, Git features. The script is not perfect but it is even better than what I was planning when I wrote my first post on the topic: A lot of checks have been put in place to make sure all variables are provided, and in the right format There is an inline help with -h The script chooses a backup randomly amongst all the available backups, not only the 30 last ones It only requires bash, cURL and jq, on a machine that has access to the Networker REST API Only tested on RHEL 6 for now The script is not complete (I already created more than 10 issues to add features, improve current features or fix things), but it is a very good start for an organisation that wants to make recovery tests easier. The Github repo will also allow...

Networker automated recovery testing using the REST API - introduction

One feature that is lacking from Networker, compared to some of its competitors, is the built-in automatic recovery testing. However, when there's an API, there is a way.  Networker's REST API is not perfect, but it allows the backup administrator to perform queries about Networker resources (objects). As my workload is going up, I realize that one of the tasks that I tend to skip the most is the periodical recovery tests.  Don't forget that a a backup that is not tested should be considered non-successful. I also found out that my recovery tests were not diverse enough. When I started this project, I knew a little bit about REST APIs, and nothing about JSON processing. With the Networker REST API documentation, and the help of a friend and Networker Support staff, I was able to create HTTP queries with Postman,  cURL and jq . Once I got the queries that I needed, I put them in a bash script that would somehow select one backup, and then restore it. My first attem...

For COVID-19, my server performance analysis course is now free!

Hi, My contribution to help people during this "stay-at-home" period: My course on Udemy is now free. You can enroll here . I hope that many of you will enroll and take the course. It's only about one hour long and have a lot of information about how to prevent, troubleshoot and solve performance problems in computing.

OTRS unmerge

Rough notes about un-merging tickets in OTRS. To delete the relationships, you simply delete relevant rows in the table named link_relation Afterwards, you re-assign articles to the right ticket (ticket_id) Regex .* Duplicated Parameter Name

MySQL backup tips

From https://hackernoon.com/elephant-in-the-room-database-backup-574da50e6d88 For instance, mysqldump will take dumps conforming to the client’s character set, and your favorite emojis such as 🍣 and 🍺 in utf8mb4 could be corrupt and replaced by  ? in the backup. If you have never checked, do it right now. Just set --default-character-set=binary option — you’re welcome. Or if you missed the --single-transaction option, you are likely to have inconsistent backups (e.g. item changed hands but money didn’t transfer) that are never easy to spot even if you regularly test the recovery procedure manually.

Postgresql tips and links

Free training material (documentation) in French: https://public.dalibo.com/exports/formation/manuels/formations/ Free book: https://books.goalkicker.com/PostgreSQLBook Documentation about the internals of PostgreSQL:  The internals of PostgreSQL  Configuration wizards https://pgconfigurator.cybertec-postgresql.com/ https://pgtune.leopard.in.ua/ https://postgresqlco.nf / is a bit different, but can help https://pgmetrics.io/   Not a wizard, but provides many info about PG instance https://postgresqlco.nf/ PG parameters documentation & info  Troubleshooting Observability: what function gives information about which component/process? https://pgstats.dev/     https://postgresqlco.nf/  You can upload your configuration, get recommendations, etc  Clients psql: PostgreSQL-provided client (command-line) pgcli CLI client with auto-completion and syntax highlighting   https://www.pgcli.com/ "Full" list Toad for PostgreSQL Kangaroo https:/...

Comware shortcuts

interface GigabitEthernet 1/0/1 => interface g1/0/1 interface Ten-GigabitEthernet 1/0/25 => interface te1/0/25 Hotkeys: <comware-switch>display hotkey ----------------- HOTKEY -----------------             =Defined hotkeys= Hotkeys Command CTRL_G  display current-configuration CTRL_L  display ip routing-table CTRL_O  undo debugging all            =Undefined hotkeys= Hotkeys Command CTRL_T  NULL CTRL_U  NULL             =System hotkeys= Hotkeys Function CTRL_A  Move the cursor to the beginning of the current line. CTRL_B  Move the cursor one character left. CTRL_C  Stop current command function. CTRL_D  Erase current character. CTRL_E  Move the cursor to the end of the current line. CTRL_F  Move the cursor one character right. CTRL_H  Erase the characte...

VIM substitutions

Substitution of a text by another text within a single line (replace I by We): :s/I/We/g Case-insensitive: :s/I/We/gi Substitute helo for hello in the next 4 lines: :s/helo/hello/g 4

VMWare vSphere 6.0 web client via SSH tunnel

Hi, I just found a way to connect to a remote vCenter server via SSH tunnel, using the web client.  This has been tested on vSphere 6.0, it may need some modifications to work on other version. Let's define some information for this example: 192.168.x.x will be the IP address of the vSphere web client (vCenter) x.x.x.x will be the IP address of the SSH server localuser is the name of the user on the local machine (from which you execute the SSH command) remoteuser is the name of the username on the (remote) SSH server 5252 is the port on which the (remote) SSH server is listening The first thing to do is to execute this shell command to open an SSH connection and create tunnels: sudo ssh -i /home/localuser/.ssh/id_rsa -l remoteuser1 -L 443:192.168.x.x:443 -L 902:192.168.x.x:902 -L 903:192.168.x.x:903 -L 9443:192.168.x.x:9443 -p 5252 x.x.x.x Please not that we must use sudo because we're using ports =< 1024. Also note the -i , specifying  the path to my pr...

Dealing with an old Rancid installation

Your new device is not supported by your current rancid install?  There is usually a solution. Get the *rancid.in and *login.in files from the github repo (https://github.com/earendilfr/rancid/tree/master/bin) Put them in the bin directory of your rancid install, rename them to remove the .in at the end chown rancid.rancid and chmod +x to these files Make sure you make the mapping in bin/rancid-fe so that the vendor is known and rancid knows which rancid script use for this vendor Edit the *rancid file to set the perl path at the top Edit the *login file to set the expect path at the top Configure your .clogin for your device Add your device to router.db Test with rancid-run -r Check the logs

General linux performance troubleshooting

Here's a list of commands that you should execute and then share the output to someone who can help you figure out what resource is the bottleneck in your system. Note, it requires the  sysstat  and  procps  packages (ubuntu and RHEL and its derivatives): uptime vmstat 1 10 iostat -xN 2 10 mpstat -P ALL 3 10 pidstat 1 10 free -mw (or free -m, if your OS doesn't support -w)  uptime is to see the load averages on the system. vmstat is mostly used to tell if the system is swapping or not. If you see significant numbers in the 'si' and 'so' columns, your system is most likely swapping (using the hard drive as RAM), which usually slows performance a lot. iostat is mostly used to determine if your disk subsystem is not able to cope with the load. If you see one or more lines that shows 100 or close almost constantly, it is probably the case.  If it is your swap volume, you probably saw numbers in the 'si' and 'so' columns in the  vmsta...

General Linux performance troubleshooting

Here's a list of commands that you should execute and then share the output to someone who can help you figure out what resource is the bottleneck in your system. Note, it requires the sysstat and  procps packages (ubuntu and RHEL and its derivatives): vmstat 1 10 iostat -xN 2 10 mpstat -P ALL 3 10 The first one is mostly used to tell if the system is swapping or not. If you see numbers in the 'si' and 'so' columns, your system is most likely swapping (using the hard drive as RAM), which usually slows performance a lot. The second one is mostly used to determine if your disk subsystem is not able to cope with the load. If you see one or more lines that shows 100 or close almost constantly, it is probably the case.  If it is your swap volume, you probably saw numbers in the 'si' and 'so' columns in the  vmstat  output. The third one shows the % of the CPU time used by different functions of the server. I'll explain the most used one...

ksar2!

After publishing my last post, I found out that many forks of kSar have been created. I didn't spend too much time comparing them, but ksar2  looks promising. Have fun!

Using kSar with Red Hat Enterprise (or CentOS) 7

Hi, I really like using kSar for troubleshooting on Linux, but it doesn't work out-of-the-box with version 7 of Red Hat Enterprise Linux or its derivatives (CentOS, Scientific Linux, etc.).  It just doesn't work with the standard " sar -a " command, so one of my colleague had a bit of free time and found the combination of arguments that makes it work. It looks like there is one of the options that are included in -a was not included in previous versions, and kSar cannot process the additional output. Here's the command line that works: sar -bBdqrRSuvwWyp -I SUM -I XALL -m ALL -n ALL -u ALL -P ALL The 'p' is, however, optional. I add it to have pretty names for disks instead of dev2-0, dev253, etc. On a side note, if you want to have better results, change the frequency at which the sar cronjob runs (in /etc/cron.d/sysstat , change the '*/10' by just '*'). Also, if you want to see 'waiting for I/O' data with RHEL 7, go in...

Another interesting package - ps_mem

I find it hard to see what's using the RAM in a server and ps_mem can help.  It lets you know the memory usage of applications. I haven't found how to install it on Ubuntu, but on Red Hat and derivatives, it's available in the EPEL repo.  You can use it by simply executing the ps_mem command, which gives this kind of output: [root@vps1 ~]# ps_mem  Private  +   Shared  =  RAM used       Program   8.0 KiB +  24.0 KiB =  32.0 KiB       agetty (2)   4.0 KiB +  30.0 KiB =  34.0 KiB       xinetd   4.0 KiB +  49.5 KiB =  53.5 KiB       lvmetad   0.0 KiB +  73.5 KiB =  73.5 KiB       saslauthd (2)   8.0 KiB +  89.0 KiB =  97.0 KiB       pmdadm  16.0 KiB + 142.5 KiB = 158.5 KiB       systemd-udevd 128.0 KiB +  90.0 KiB =...

ncdu (NCurses Disk Usage)

Image
I discovered an utility a few weeks ago and I thought it would be worth sharing. This utility is especially useful when you are looking for files that take a lot of disk space on your system.  As an example, I received a Nagios notification saying that /var had less than 15% free.  I logged on the system to see what is taking too much space. In the past, I would have use the du command (du -hs /var/*, then du -hs /var/log/* and so on), but I used ncdu.  I cd'd into /var, then ran 'ncdu'. Here are screen shots of an example (different of what happened yesterday): This is what we see when we start ncdu in /var: We can use the arrows up and down to select a folder and use the left and right arrows to go up/down the tree. In this case, I simply used the right-hand side arrow a few times until I got to the folder that contains the most data: That's already great, but that's before digging into the help (using ?).  The options are (version 1.13): Sortin...