Red Hat es mundialmente conocida por sus certificaciones, ya que sus exámenes distan mucho de las tradicionales pruebas de opción múltiple o incluso de los exámenes con preguntas abiertas, pues en Red Hat, los exámenes son totalmente prácticos, te ponen delante de una máquina, con una lista de actividades que se tienen que realizar sin conexión a internet, siendo monitoreado remotamente por dos cámaras web por el aplicador y teniendo pocos descansos (sin que estos detengan el tiempo disponible para realizar el intento de certificación). Por estas y otras razones (como el costo de sus cursos e intentos de examen), el peso específico que tienen estas certificaciones en la industria de las TI es bastante alto.
Hace poco comencé a recorrer el certification path que la empresa ofrece para llegar a ser algún día un Red Hat Certified Architect (RHCA), camino que requiere dos certificaciones base que son la Red Hat Certified System Administrator (RHCSA) y la Red Hat Certified Engineer (RHCE) más otras cinco certificaciones que pueden ser de muchas tecnologías, dependiendo más de gustos y necesidades laborales; así pues, en total son siete certificaciones para lograr ser un RHCA.
Pues bien, dado que ya tengo la RHCSA, hablaré de mi plan de ataque para obtener la RHCE al primer o segundo intento, pues, esta certificación tiene fama de ser muy pesada, poca gente reporta terminar todas las actividades en el tiempo disponible (4 horas) así como su nivel de dificultad, pues como dije, uno está solo contra la máquina, pudiéndose consultar solamente la documentación del sistema operativo y además siendo monitoreado todo el tiempo.
El temario básicamente es el siguiente, obviamente todo se basa en Red Hat Enterprise Linux 7 hasta el momento, aunque ya no tarda RHEL 8:
Configuración de servicios con Systemd
Control del proceso de arranque
Configurar redes IPv4
Configurar redes IPv6
Configuración Link Aggregation
Configuración de bridges por software
Administración de Firewalld
Etiquetado de puertos con SELinux
Configurar servidor DNS (unbound)
Solucionar problemas de DNS
Configurar servidor de correo
Conceptos iSCSI
Configurar servidor NFS
Configurar servidor SMB/CIFS
Administración e instalación de María DB
Queries RDBMS MySQL/Maria DB
Configurar servidor HTTPD Apache
Hosts virtuales y aplicaciones web con Apache
Shell scripting con Bash
Shell scripting avanzado, estructuras de control y condiciones
Uso de Docker containers
Realmente no son tópicos del otro mundo, son cosas que uno como sysadmin ha tenido que realizar al menos una vez durante la experiencia laboral, sin embargo, es menester estar bien preparado, con los comandos y los archivos de configuración frescos en la memoria y con la capacidad de solucionar los problemas más comunes de forma rápida porque el tiempo avanza velozmente y cuando menos te das cuenta, te quedan 30 minutos para terminar de responder.
Espero en entradas próximas ir desmenuzando este temario para que sea útil a más personas.
Actualmente cuento ya con cuatro certificaciones de Red Hat, a saber:
Sin duda, vivimos en tiempos donde la dinámica económica está cambiando en favor de los Servicios, donde la llamada Economía del Conocimiento es cada vez más palpable y está llegando a cada vez más personas en todo el mundo. Vivir en esta nueva realidad puede ser muy ventajoso, ya que por primera vez, se tiene un potencial de ascenso social no condicionado por el origen de la persona, sino por su talento y sus conocimientos.
Mucho se habla de cómo la educación formal, básica, media y superior han ayudado a millones de personas a mejorar su nivel de vida y esto es cierto, pero también hay que mencionar que esta educación no es suficiente, pues vivimos una tendencia de especialización de los puestos de trabajo, en detrimento de labores aburridas, repetitivas y poco gratificantes, lo cual, dicho sea de paso, me parece algo muy bueno.
En el mundo de la computación, he de decir, que las clases formales siempre me parecieron (y me parecen) muy aburridas, incluso las que trataban de los temas que más me gustaban, con excepción de las clases realmente teóricas (y donde la cátedra tradicional no puede ser sustituida), como estructuras de datos, cálculo, álgebra superior, matemáticas discretas, etc. Pero en el caso de las materias más prácticas, como las de programación, simplemente las encontré muy tediosas y creo que era por una simple razón: ver láminas o presentaciones con porciones de código sin realmente ejecutarlo uno mismo es como ver ruido de las televisiones de antaño, ruido que es descartado por el cerebro casi de inmediato. Claro, hasta el momento que intentas programar algo tú mismo, no sabes cómo hacerlo y recuerdas vagamente al profesor intentando enseñarte justamente eso.
Entonces,
¿Cómo adquirir estos conocimientos que ayudan a marcar la diferencia de una mejor forma?
No hay secreto: haciéndolo uno mismo.
Aquí es donde entran portales como Katacoda, que proporcionan plataformas bastante interesantes, donde uno puede ir practicando directamente en una consola y en interfaces web embebidas, de modo que ya no hay pretexto para no aprender algo nuevo y de una forma divertida (para los frikis jaja)
Página de inicio de Katacoda
En particular, esta herramienta contiene muchos escenarios que sirven para ir aprendiendo sobre las nuevas tecnologías como Docker, Kubernetes, Jenkins, OpenShift, NodeJS, Cloud Platforms y hasta un poco de machine learning. Por supuesto, una gran parte del contenido está disponible de forma gratuita, esperando a que alguien quiera aprender haciendo.
Well, this is a very technical post and, in order to be useful for more people, I decided to write it in English.
Prerequisites
First of all, you should be able to have access to the OpenShift Enterprise repositories from Red Hat. Disclaimer: the current post is not related to CentOS based installations, also this post is not referring to OKD, MiniShift or even OpenShift on a single container (all-in-one), nope, this post is the technical review of OpenShift Enterprise 3.11 installed on a relatively small hardware.
Hardware
Intel NUC Core i7 quad core, 32 GB of RAM, 256 GB PCIe Flash, 750 GB Hard Disk.
Intel NUC Core i5 quad core, 16 GB of RAM, 256 GB PCIe Flash, 500 GB Hard Disk.
Intel NUC Core i3 quad core, 16 GB of RAM, 128 GB PCIe Flash, 500 GB Hard Disk.
HP MP 9, Intel Core i5 quad core, 16 GB of RAM, 256 SSD. This is the master node.
Lenovo Laptop with RHEL 7.6 as bastion host.
The nodes of my cluster.
Cluster components
1 Master node
3 Infra-compute nodes (infrastructure and computing nodes)
3 GlusterFS nodes
As you can advice, there are not enough nodes to cover the proposed architecture, so, I'm proposing to share the node resources in order to deploy the infra-compute nodes and the glusterfs nodes together. Of course, this kind of deployment is not recommended by Red Hat, remember, this is only a cluster for learning purposes, never for production purposes.
Brief list of OpenShift services to deploy
Hawkular
Cassandra
Heapster
Elasticsearch
Fluentd
Kibana
Alert manager
Prometheus
Grafana
GlusterFS
Web Console
Catalog
Cluster console
Docker registry
OLM Operators
Problem detector
OC command line
Heketi
Master API
Internal Router
Scheduler
and more...
Operating system
Red Hat Enterprise Linux (RHEL) 7.6 up to date with the following repositories enabled:
The required RPMs on every node and the bastion host should be: wget, git, net-tools, bind-utils, yum-utils, firewalld, java-1.8.0-openjdk, bridge-utils, bash-completion, kexec-tools, sos, psacct, openshift-ansible, glusterfs-fuse, docker and skopeo.
For the GlusterFS nodes (in this case, the nodei7, nodei5 and nodei3 nodes, also install Heketi Server on the master node) you should to install also the Heketi packages and the GlusterFS server packages:
Adding these entries on the /etc/hosts file is not sufficient to deploy successfully the cluster due to OpenShift generates automatically internal rules on the Kubernetes pods and it takes them from the DNS configuration.
master.calvarado04.com
nodei7.calvarado04.com
nodei5.calvarado04.com
nodei3.calvarado04.com
*.openshift.calvarado04.com <--- Wildcard, it must be making reference to the nodes where is deployed the router service (infra nodes).
cluster-openshift.calvarado04.com <--- Master node public name
Don't forget to change the name of your nodes in the same way, you can perform that by:
SSL Certificate (wildcard capable) for the OpenShift cluster
As you can see, If you want to install your own SSL certificate, that certificate must have wildcard capabilities, due to the basic SSL certificates won't work at all on OpenShift.
Your common name on your CSR request should be like this for the router certificate:
*.calvarado04.com for the master
*.openshift.calvarado04.com for the router subdomain (all the apps)
I added my certificate files on the bastion host on the directory /root/calvarado04.com, the files needed are:
/root/calvarado04.com/calvarado04.pem (that is the certificate file merged with the intermediate certificate and the root certificate, the certificate content at the top of the file).
/root/calvarado04.com/calvarado04.key with the private key generated when I generated the CSR file to obtain the SSL certificate.
/root/calvarado04.com/calvarado04.ca with the root certificate.
And the same for the router certificates (Openshift subdomain) on /root/calvarado04.com/openshift/
SELinux
Many people just disable SELinux on almost all new installation, this is not the case, SELinux is mandatory, it must be targered and enforcing on all nodes on the /etc/selinux/config file.
# This file controls the state of SELinux on the system.
# SELINUX= can take one of these three values:
# enforcing - SELinux security policy is enforced.
# permissive - SELinux prints warnings instead of enforcing.
# disabled - No SELinux policy is loaded.
SELINUX=enforcing
# SELINUXTYPE= can take one of these three values:
# targeted - Targeted processes are protected,
# minimum - Modification of targeted policy. Only selected processes are protected.
# mls - Multi Level Security protection.
SELINUXTYPE=targeted
Also please add the following rules to give permissions to the containers:
[root@bastion ~]# ansible nodes -a "setsebool -P virt_sandbox_use_fusefs on"
[root@bastion ~]# ansible nodes -a "setsebool -P virt_use_fusefs on"
Firewalld
OpenShift used to work with IPtables, but this is not recommended anymore and any new deployment should be using firewalld instead. To mask and disable iptables:
OpenShift uses the NetworkManager capabilities, so, is required to have enabled NetworkManager on the network device, for instance, check the following configuration:
Ansible requires SSH keys to run the required tasks on each node, let's create a key and distribute it to every nodes:
[root@bastion ~]# ssh-keygen
[root@bastion ~]# for host in master.calvarado04.com \
nodei7.calvarado04.com \
nodei5.calvarado04.com \
nodei3.calvarado04.com; \
do ssh-copy-id -i ~/.ssh/id_rsa.pub $host; \
done
Ansible
The official OpenShift Enterprise 3.11.69 installation method is by running the OpenShift playbooks provided by the RPMs. The Ansible version must be 2.6.x not 2.7 due to issues with some tasks.
GlusterFS
This deployment will be using GlusterFS on the most basic installation using only GlusterFS without blocks and is not including the GlusterFS Registry cluster on Convergent mode. However, this method is allowing you to use dynamic provisioning of the PVC's (Persistent Volume Claims), this feature is quite fancy because is not longer needed to declare manually the PV's once the cluster is up and running.
Heketi configuration on the master node
OpenShift uses Heketi to generate the Gluster topology from the master node to the Gluster nodes and it perform that task by using a SSH connection. That's why you need to configure Heketi on the master host by adding the user and ssh key on /etc/heketi/heketi.json.
Also change any timeout entry to a numeric value (on _sshexec_comment and _kubeexec_comment), like "gluster_cli_timeout": 900.
On my first attempt to deploy the cluster I was stuck on a step that were wait for Gluster pods, the counter was timed out and the playbook was marked as failure. I checked the pods on the nodes and I saw them, I checked its logs and its logs was not showing any error. Researching more I found that could be a bug on the wait_for_pods.yml deployment playbook, located on /usr/share/ansible/openshift-ansible/roles/openshift_storage_glusterfs/tasks, concretely a cast issue due to the playbook is comparing a string with an integer.
This is the original playbook:
---
- name: Wait for GlusterFS pods
oc_obj:
namespace: "{{ glusterfs_namespace }}"
kind: pod
state: list
selector: "glusterfs={{ glusterfs_name }}-pod"
register: glusterfs_pods_wait
until:
- "glusterfs_pods_wait.results.results[0]['items'] | count > 0"
# There must be as many pods with 'Ready' staus True as there are nodes expecting those pods
- "glusterfs_pods_wait.results.results[0]['items'] | lib_utils_oo_collect(attribute='status.conditions') | lib_utils_oo_collect(attribute='status', filters={'type': 'Ready'}) | map('bool') | select | list | count == l_glusterfs_count | int"
delay: 30
retries: "{{ (glusterfs_timeout | int / 10) | int }}"
vars:
l_glusterfs_count: "{{ glusterfs_count | default(glusterfs_nodes | count ) }}"
And this is the fixed playbook:
---
- name: Wait for GlusterFS pods
oc_obj:
namespace: "{{ glusterfs_namespace }}"
kind: pod
state: list
selector: "glusterfs={{ glusterfs_name }}-pod"
register: glusterfs_pods_wait
until:
- "glusterfs_pods_wait.results.results[0]['items'] | count > 0"
# There must be as many pods with 'Ready' staus True as there are nodes expecting those pods
- "glusterfs_pods_wait.results.results[0]['items'] | lib_utils_oo_collect(attribute='status.conditions') | lib_utils_oo_collect(attribute='status', filters={'type': 'Ready'}) | map('bool') | select | list | count == l_glusterfs_count | int"
delay: 30
retries: "{{ (glusterfs_timeout | int / 10) | int }}"
vars:
l_glusterfs_count: "{{ glusterfs_count | default(glusterfs_nodes | count ) }} | int"
Wipe the GlusterFS disks
In order to install successfully the GlusterFS cluster, please wipe all the data from the chosen disks (in this case, the three GlusterFS nodes have the disk on /dev/sda, but that it could be different in your configuration, be careful), there must be on RAW format only. Warning: you can lose all your data if you are not careful. You can perform this by:
[root@bastion ~]# ansible glusterfs -a "wipefs -a -f /dev/sda"
HTPasswd (users and passwords for OpenShift)
This is a really simple configuration on this topic, I'm using only htpasswd, I saved my users on /root/htpasswd.openshift on the bastion host. You should to try with LDAP, it's more secure and advanced. Go ahead and learn something 😉
If you want to use the same way just:
[root@bastion ~]# htpasswd -c /root/htpasswd.openshift luke
New password:
Re-type new password:
Adding password for user luke
[root@broker ~]# cat /root/htpasswd.openshift
luke:$apr1$SzuTxyhH$DlH976Tv2cDBccFZqJ3zf1
If you want to add more users, just omit the -c option:
[root@bastion ~]# htpasswd /root/htpasswd.openshift leia
New password:
Re-type new password:
Adding password for user leia
[root@bastion ~]# cat /root/htpasswd.openshift
luke:$apr1$SzuTxyhH$DlH976Tv2cDBccFZqJ3zf1
leia:$apr1$UK2B49gy$qD/n3lKsoWXT0eRAMNoqm.
Run the installation playbooks
Once the package openshift-ansible is already installed on your bastion host and your SSH keys are already distributed on all your nodes, you can perform the installation of your OpenShift Cluster.
The happy path is quite simple:
Fill your inventory file located in /etc/ansible/hosts