Monday, October 30, 2023

Proxmox VE Cluster - Chapter 004 - Infrastructure Prep

Proxmox VE Cluster - Chapter 004 - Infrastructure Prep


A voyage of adventure, moving a diverse workload running on OpenStack, Harvester, and RKE2 K8S clusters over to a Proxmox VE cluster.


Before I change anything, I wanted to prepare as much infrastructure as possible.  I did not want to interrupt a conversion step via the realization I forgot to allocate IP addresses or forgot to configure the DNS server.


For day-to-day documentation, such as lists of IP address or URL links to services, I use Dokuwiki which is the most basic and simple wiki software I can find.  I created a new wiki page and added the entire IP allocation for the new cluster:

  • proxmox001 (old os1 node in OS1 cluster) SuperMicro SYS-E200-8D 10.10.8.1
  • proxmox002 (old os2 node in OS1 cluster) SuperMicro SYS-E200-8D 10.10.8.2
  • proxmox003 (old os3 node in OS1 cluster) SuperMicro SYS-E200-8D 10.10.8.3
  • proxmox004 (old os4 node in OS2 cluster) SuperMicro SYS-E200-8D 10.10.8.4
  • proxmox005 (old os5 node in OS2 cluster) SuperMicro SYS-E200-8D 10.10.8.5
  • proxmox006 (old os6 node in OS2 cluster) SuperMicro SYS-E200-8D 10.10.8.6
  • proxmox011 (old harvester-small-1) Intel NUC6i3SYH 10.10.8.11
  • proxmox012 (old harvester-small-2) Intel NUC6i3SYH 10.10.8.12
  • proxmox013 (old harvester-small-3) Intel NUC6i3SYH 10.10.8.13
  • proxmox021 (old rancher1) Beelink N5095 10.10.8.21
  • proxmox022 (old rancher2) Beelink N5095 10.10.8.22
  • proxmox023 (old rancher3) Beelink N5095 10.10.8.23
  • proxmox031 (old docker) Intel NUC6i3SYH 10.10.8.31
  • proxmoxbackup (old server bare metal hardware) 10.10.8.254
  • freenas.cedar.mulhollon.com 10.10.20.4 (use IP addresses for NFS so as to not rely on DNS)
As systems came online, I added hyperlinks to their web GUI and generally keep the wiki page up-to-date with the current configuration.  If I'm working on the cluster, I probably have this wiki page open.


Set up Redmine project for Proxmox:

I use Redmine for project level long term, large scale documentation.  I created a project in Redmine to basically be a more detailed version of the wiki page.


Set up simple insecure NFS on TrueNAS for non-production short term testing:

https://pve.proxmox.com/wiki/Storage:_NFS

https://pve.proxmox.com/pve-docs/chapter-pvesm.html#storage_nfs

A list of six content types to create shares for in TrueNAS (as documented in the Wiki and in Redmine, of course):

  1. proxmox-containers
  2. proxmox-containertemplates
  3. proxmox-diskimages
  4. proxmox-isoimages
  5. proxmox-snippets
  6. proxmox-vzdumps

Creating the six NFS exports in TrueNAS for Promox:

  1. First create the datasets, for later NFS export.  
  2. "Storage", "Pools", in freenas-pool, three dots "Add Dataset".
  3. Name: the id from the list
  4. Comments: "something that makes sense"
  5. Compression Level: off
  6. "Submit"
  7. Then export the six datasets.
  8. "Sharing", "Unix Shares (NFS)",  "Add", "Advanced Options"
  9. Path: /mnt/freenas-pool/proxmox-diskimages (or similar)
  10. Maproot User blank
  11. Maproot Group blank
  12. Mapall User: root
  13. Mapall Group: wheel
  14. Optionally add authorized networks and hosts, later.  No need to access outside 10.0.0.0/8, obviously.
  15. Add symlinks in my homedir to the automounter locations "ln -s /net/freenas/mnt/freenas-pool/proxmox-diskimages ~"
  16. Test by creating and deleting some files.


Add sysadmin issue tasks in Redmine in the "Systems Administration" project for all cluster nodes.


Document all the changes and allocations in Netbox

Here is the link to the Port / Services list for Proxmox VE nodes:

https://pve.proxmox.com/pve-docs/chapter-pvecm.html#_requirements


Set up a USB flash boot drive for servers that can't PXEboot

Download Proxmox VE 8.0-2 and put it on a properly labeled USB flash drive.

https://www.proxmox.com/en/downloads/proxmox-virtual-environment/iso

https://pve.proxmox.com/pve-docs/chapter-pve-installation.html#installation_prepare_media


At this point I think I've prepared everything possible.  In the next post, I start the conversion work.

Friday, October 27, 2023

Proxmox VE Cluster - Chapter 003 - Order of Operations for Level 1.0

Proxmox VE Cluster - Chapter 003 - Order of Operations for Level 1.0


A voyage of adventure, moving a diverse workload running on OpenStack, Harvester, and RKE2 K8S clusters over to a Proxmox VE cluster.


The order of operations is complicated because I intend to keep most of the workload fully operational during the conversion.  Its kind of like rebuilding a car engine while driving the car down the road.


  1. Convert the old Rancher "RKE2" K8S cluster into a very small Proxmox VE cluster.  This step will be successful if I have a working three node Proxmox cluster.
  2. Move everything off the small Rancher Harvester test cluster, which is currently slowly running an old version of Harvester, and add those nodes into the small Proxmox VE cluster.  The measure of success will be having six clustered Proxmox nodes.
  3. Get some practice with VMs on the small test cluster.  Success will be defined as some working scratch (test load only) Ubuntu servers with a couple days "burn-in" and operational experimentation.
  4. Move a minimal set of very small production services on the small Proxmox VE cluster.  Maybe start with the wiki server, one of the multiple Active Directory Domain Controllers, just enough to prove out the operation of the cluster.  I consider this step a success if everything "important" on the old OS1 cluster is minimally running on the new Proxmox server reliably for a couple days.
  5. Migrate all production workload off the OS1 OpenStack cluster then add the former OS1 nodes into the now medium-sized Proxmox VE cluster.  Success at this step looks like a Proxmox cluster of nine nodes running a minimal production workload for a couple uninterrupted days.
  6. Roll ALL production workload off the OS2 OpenStack cluster into the now medium sized Proxmox VE cluster, likely a tight fit.  Success looks like OS2 having zero load and Proxmox carrying the entire production load, although in theory if Proxmox crashed I have the OS2 cluster has a hot-backup to Proxmox.
  7. Convert the remaining OS2 OpenStack cluster into even more Proxmox cluster capacity.  This step is a success if the Proxmox cluster has twelve operating nodes holding the entire production load, and I'm no longer running RKE2 or Harvester or OpenStack or any other cluster system on bare metal.
  8. Verify Operation, load balancing, run it for awhile before working on Architecture level 2.0.  Success looks like no crashing, no bugs, no issues, optimized CPU/memory settings and optimized workload across the dozen cluster nodes.

Next post will be about Infrastructure preparation efforts, get as much stuff ready as possible before starting the big conversion project.

Wednesday, October 25, 2023

Proxmox VE Cluster - Chapter 002 - Plan for Architecture Level 1.0

Proxmox VE Cluster - Chapter 002 - Plan for Architecture Level 1.0


A voyage of adventure, moving a diverse workload running on OpenStack, Harvester, and RKE2 K8S clusters over to a Proxmox VE cluster.


Here's my detailed plan for Architecture Level 1.0:


In summary, Level 1.0 means drop everything from multiple old clusters into a single large, simple-as-possible Proxmox VE cluster via several conversion phases, while making absolute minimum changes to design and workload.


A list of Goals for Architecture Level 1.0: 

  1. All production storage on the giant TrueNAS over NFS.  Cluster-wide filesystems implemented later.
  2. VMs manually provisioned, like the older VMWare era.  I like Terraform and Ansible as tools to provision IaaS, but I will implement that later.
  3. Generally document and make minimal changes in Ansible with respect to the workload VMs and containers.


Some of the eternal ongoing projects such as re-IP addressing will continue as part of Arch 1.0, which is ambitious.  In a way it makes the conversion from OpenStack and Harvester to Proxmox VE simpler, if the new VM has a new IP address.  As usual I will polish and refine my Netbox information, clean up runbooks stored in Redmine, but the theme will be making minimum-possible changes rather than implementing ambitious new ideas at the same time as the cluster conversion.


The question of Docker Containers...


https://pve.proxmox.com/wiki/Linux_Container

"If you want to run application containers, for example, Docker images, it is recommended that you run them inside a Proxmox QEMU VM. This will give you all the advantages of application containerization, while also providing the benefits that VMs offer, such as strong isolation from the host and the ability to live-migrate, which otherwise isn’t possible with containers."


OpenStack Zen containers were cool, and K8S obviously runs container workloads very well, however the above Proxmox information implies I will have to go back to the VMware era of setting up "container containers" to hold my Docker containers.  No big deal, but it will take some time to roll everything.  Generally I do "container containers" by installing a simple Ubuntu server, then running Docker off a NFS mount so there is no local state stored on the Ubuntu server (making it trivial to rebuild, and also making backups very simple as all state is just files on the NFS server).


In the next post, I will discuss the complicated order of operations due to various dependencies and operations requirements.

Monday, October 23, 2023

Proxmox VE Cluster - Chapter 001 - Why switch to Proxmox VE?

Proxmox VE Cluster - Chapter 001 - Why switch to Proxmox VE?


A voyage of adventure, moving a diverse workload running on OpenStack, Harvester, and RKE2 K8S clusters over to a Proxmox VE cluster.


Why switch from OpenStack / Harvester / RKE2 K8S, VMware, and other tech to a standard data center baseline of Proxmox VE?

  • The hardware load and requirement is high for Harvester.  Harvester is awesome but some of my smaller nodes spend most of their CPU cycles and memory running Harvester itself rather than running my workloads.  A full Rancher cluster to control Harvester is very cool technology, but the hardware load is expensive.
  • Harvester upgrades fail because the hardware load is too high.  Related to the above, I'm having trouble upgrading the more heavily loaded nodes because they can barely run Harvester at all, much less afford the extra system load to upgrade K8S.
  • I can't really go backward to VMware.  Hardware compatibility lists, etc.  Technically I could spend the money but I don't think it would be worth it.
  • I've learned everything I can learn from OpenStack and looking at trends it's time to 'jump ship' from OpenStack.
  • Upgrades for OpenStack kolla-ansible are non-trivial and a little more interactive than I would prefer.  I'm running two OS clusters and will push the workload over to one cluster temporarily while upgrading the other cluster.  It takes a lot of time and sysadmin effort.
  • Proxmox has some cool new features to experiment with, like native CEPH distributed cluster filesystem integrated into the system, and the very cool looking Proxmox Backup Server system.
  • I want a "single pane of glass" to manage my cluster hardware with respect to monitoring, control, backup, etc.  I will put everything on Proxmox and control everything via a single Proxmox cluster.  I don't want a RKE2 K8S cluster AND a Harvester cluster AND two OpenStack clusters to manage, just one big Proxmox cluster, ideally.


Very high level plan for the overall conversion project:

  1. Architecture level 1.0 will be a phased conversion of all workload into a large Proxmox VE cluster.  This will be an OpenStack / Harvester type design and workload wedged into fitting in to Proxmox VE.
  2. Architecture level 2.0 will be system integration to make this a Proxmox-styled cluster, integrating with monitoring and automation and generally adapting the workload to make everything feel integrated rather than a different system's workload being temporarily run on Proxmox.
  3. Architecture level 3.0 will be more R+D focused, advancing into interesting extra features that are Proxmox-specific, such as a cluster wide filesystem, the Proxmox Backup Server system, some interesting advanced networking ideas, new stuff in general.


How did I get here?

Can't figure out how to get where you're going, unless you know how you got where you are right now.

  • Around the turn of the century, had bare metal Linux servers running LXC and also the FreeBSD equivalent.
  • Around the 2010s, had some sysadmin level experience with VMware at work, and I signed up for the "ESXi Evaluation Experience" which gives you limited non-commercial license to pretty much the entire collection of VMware software.  This was pretty awesome for several years, although the continual drift in the hardware compatibility list and hardware requirements increasing dramatically over time, and generally being tired of paying for even a discounted VMware license, meant I moved away from VMware.
  • Around the late 2010s / 2020 timeframe, replaced the VMware cluster with OpenStack.  OpenStack is FOSS, but the labor required to keep it up is expensive.
  • In the early 2020s I started experimenting with RKE2 K8S on bare metal, and Rancher's Harvester bare metal virtualization solution.  Nice tech and works well, but it's designed for "larger" individual nodes than I can afford.

The above leads me to being interested in Proxmox VE to underlie my entire mini-datacenter.  I will eventually put everything on Proxmox except for my NAS and a stand alone Proxmox backup server.

Something to note about this series is the blog posts appear some time after "the action".  So as you read a new blog post, this all happened some weeks / months ago.  This gives me time to document things I've missed, circle back around, etc.  On the bad side I might miss some details of something I did months ago.  On the good side, I've circled back to document bugs and workarounds and any other areas of friction, which should save you, as the reader, some time if you implement a Proxmox cluster at your site.

Next post in the series will be a more detailed description of my Architecture level 1.0.

Friday, March 10, 2023

Rancher Suite K8S Adventure - Chapter 020 - Prepare Terraform for Harvester

Rancher Suite K8S Adventure - Chapter 020 - Prepare Terraform for Harvester

A travelogue of converting from OpenStack to Suse's Rancher Suite for K8S including RKE2, Harvester, kubectl, helm.

I don't like manually configuring things.  I like IaaC with templates stored in a nice Git repo, it eliminates errors, deployments are faster, fewer human errors, it's just all around better than mousing and typing a virtual infrastructure.  So today we prepare Terraform to work with Harvester, but first, some work with multiple cluster kubeconfig files.

Multiple Cluster Kubectl

Start automating by configuring kubectl to talk to multiple clusters.

Reference:

https://kubernetes.io/docs/tasks/access-application-cluster/configure-access-multiple-clusters/

The good news is its pretty easy to configure multiple clusters into separate kubectl contexts.  The bad news is its easy to select different contexts at runtime, in fact its so easy to select different contexts, that there have been multiple headline news stories about devops who thought they were permanently erasing their test cluster deployment, only to rapidly discovery they were actually in their production context, resulting in some amazing news headlines about outages and deleted data.  So, keep your wits about you and be careful.  I will set up multiple contexts some other time.

One of the cultural oddities of the K8S community is they like to call the kubectl config file by the generic phrase "your kubeconfig file".  What makes that odd is most installs do not have a file named kubeconfig or dot kubeconfig or kubeconfig.conf or whatever.  On my Ubuntu system, kubectl's config file, aka the "kubeconfig file" is configured by a file located at ~/.kube/config

I will usually be working with Harvester, so in my ~/.kube directory I keep yaml files named rancher.yaml and harvester.yaml and I can simply copy them over the ~/.kube/config file.

In summary, make certain that running "kubectl get nodes" displays the correct cluster... 

Terraform

https://developer.hashicorp.com/terraform

Terraform is similar in concept to CloudFormation from AWS or HEAT templates from OpenStack.  You write your infrastructure as source code, run the template, and terraform makes the cloud gradually closely resemble your template.  Not a script, so much as a specification.

Install Terraform

I should have installed Terraform back when I was installing support software like kubectl and helm.  Better late than never...

https://developer.hashicorp.com/terraform/downloads

https://www.hashicorp.com/official-packaging-guide

The exact version of the Ubuntu package I'm installing is 1.3.9 as seen at

https://releases.hashicorp.com/terraform/

And I'm doing an "apt hold" on it to make sure its not accidentally upgraded.

Here is a link to the Gitlab repo directory for the Ansible helm role:

https://gitlab.com/SpringCitySolutionsLLC/ansible/-/tree/master/roles/terraform

If you look at the Ansible task named packages.yml, the task installs some boring required packages first, then deletes the repo key if its too old, then downloads a new copy of the repo key if its not already present, gpg dearmor the key into 'apt' format, add the local copy of the repo key to apt's list of known good keys, install the sources.list file for the repo, does an apt-get update, takes terraform out of "hold" state, installs terraform version 1.3.9, finally places terraform back on "hold" state so its not magically upgraded to the latest version (1.4 or 1.5 or something by now).  Glad I don't have to do that manually by hand on every machine on my LAN, LOL.

Simply add "- terraform" to a machine's Ansible playbook, then run "ansible-playbook --tags terraform playbooks/someHostname.yml" and it works.  Ansible is super cool!

As of the time this blog was written, "terraform --version" looks like this:
vince@ubuntu:~$ terraform --version
Terraform v1.3.9
on linux_amd64
vince@ubuntu:~$ 

References

https://www.suse.com/c/rancher_blog/managing-harvester-with-terraform/

https://docs.harvesterhci.io/v1.1/terraform/

https://github.com/harvester/terraform-provider-harvester

https://registry.terraform.io/providers/harvester/harvester/latest


Thursday, March 9, 2023

Rancher Suite K8S Adventure - Chapter 019 - Provision a VM in Harvester

Rancher Suite K8S Adventure - Chapter 019 - Provision a VM in Harvester

A travelogue of converting from OpenStack to Suse's Rancher Suite for K8S including RKE2, Harvester, kubectl, helm.

A large proportion of Blog / Youtube videos covering Rancher seem to halt after the UI is working.  I intend to extend from that point, and document using the system.  Today we provision a VM.  Personally I do the hybrid life, I provision a VM in AWS, VMware, OpenStack, or now Harvester, then I use Ansible to automatically integrate the VM into my existing systems WRT Active Directory SSO, Logging, Zabbix and MetricBeats monitoring, NTP, etc.  As per that operational style, the goal or demarcation point of provisioning a VM would be successfully SSH in as the Ansible user, beyond that point Ansible takes over configuration.

As a general observation over the years the workload has steadily become container-ized such that in the long run I don't think I'll have many non-infrastructure VMs.  I will likely continue to have multiple VMs for DCs, DNS, DHCP, maybe a few other tasks, but workload slowly always moves toward containers.  That should work very well with a HCI solution such as Harvester.  Use case drives the requirements; for example I will have to bridge the DHCP server interface directly onto the LAN, for example, which was "easy" in OpenStack.

Rancher Project vs K8S Namespace

https://ranchermanager.docs.rancher.com/pages-for-subheaders/manage-projects

https://ranchermanager.docs.rancher.com/how-to-guides/new-user-guides/manage-namespaces

Projects are a 'new' Rancher concept wedged between existing K8S clusters and existing K8S namespaces.  Clusters contain projects which contain namespaces.  My plan is to use projects in Rancher similar to how I used projects in OpenStack, so I will configure "infrastructure" "server" "iot" "enduser" and similar project names.

Note that projects are a part of a cluster; project "infrastructure" on cluster "harvester-small" is independent of project "infrastructure" on cluster "harvester-large".

I intend to create roughly one namespace per hostname or system as appropriate.  Everything about DHCP as a system, would live in the DHCP namespace in the infrastructure project in my Harvester cluster.

Create a project

Log into Rancher, select virtualization and the harvester-small cluster, hamburger menu "Projects/Namespaces", "Create Project"

I named today's test project "experiment", and added "admin" as a project owner, so there are two owners, user "vince" (me) and the "admin" user.

This is where you would set project resource quotas and limits, if you were planning to use any (which I am not, at this time)

Create a namespace

Log into Rancher, select virtualization and the harvester-small cluster, hamburger menu "Projects/Namespaces", in the "experiment" project area click "Create Namespace".

I named my namespace "vm".  This is where you can set container level resource limits.  I don't intend at this time to set any limits to this namespace.

Create a VM storage class

I don't need 3 replicas of a test volume on a 3 node cluster, 2 replicas should be fine for my testing.

https://docs.harvesterhci.io/v1.1/advanced/storageclass

Log into Rancher, select virtualization and the harvester-small cluster, click "Advanced", then "Storage Classes".

Note the default SC is "harvester-longhorn" and it keeps 3 replicas.

Click 3-dots, "clone", give it a name "experimental", change the number of replicas to 2, click "create".

Upload an image into the cluster

https://docs.harvesterhci.io/v1.1/upload-image

Let's use Ubuntu's 20.04 focal-server-cloudimg-amd64.img file, I'm using version 20230215 from:

http://cloud-images.ubuntu.com/focal/20230215/

Note this is exactly the same image I run on OpenStack.  The similar named "-disk-kvm" optional image has a console compatibility issue with OpenStack such that if you want a working console you can't use "-disk-kvm".  For consistencies sake I will use the same image on Harvester.  This console works fine in OpenStack and in Harvester.

First we try uploading a previously downloaded cloud image.  Spoiler alert, this does not work.  Log into Rancher, select virtualization and the harvester-small cluster, click "Images" then "Create".  I'm naming my upload "Ubuntu 20.04 20230215".  Then select the img file and upload it.  It will take awhile to leave "uploading" status.  Eventually it failed with "Timeout waiting for the datasource file processing begin".  Well OK, let's retry.  That fails again.  There is an old Harvester bug report in Github about this error message that is closed, but obviously file uploads still don't work:

https://github.com/harvester/harvester/issues/1415

OK then, I will upload via providing a URL instead of download/upload.  "Copy Link Address" from the Ubuntu cloud-images page, and do a URL upload directly into Harvester.  The image download via URL succeeded after about three minutes.  Cool.

For now, I am using "Storage" "Storage Class" the default harvester-longhorn, which makes triple copies of the image, which seems excessive for an internet download; on the other hand I think maybe VMs will instantiate faster if the copy is local to the hardware.  I will experiment later on, with creating an image storage class that only keeps one replica to save disk space.  After all, I only need read access at the moment of instantiating a new VM which is pretty rare, and if I lose the image I can re-download a new copy from the internet in less than three minutes...

Upload your SSH keys

https://docs.harvesterhci.io/v1.1/vm/create-vm

I have a regular daily use SSO Active Directory user with SSH keys AND an ansible user with SSH keys, so for experiments (like today) I would provision a VM with my "vince" user SSH keys but I would provision a real VM with the "ansible" user SSH keys.  So I have two sets of keys to upload although today we're only using the "vince" ssh key.

Log into Rancher, select virtualization and the harvester-small cluster, click "Advanced" then "SSH Keys" then "Create".  Note that you can put SSH keys into specific namespaces although I'm using "default".  "Read from a File" worked for me.

Repeat the above for the "ansible" user, which we won't be using today but will use sooner or later.

Create a VM Network

By default, if you try to create a VM, the VM will be placed in network "management Network" which is the internal system network, and you'll get some crazy inaccessible 10.50.x.y address that only exists inside the cluster.  So you need to configure a "VM Network" that bridges over to the ethernet port, connecting the VM to the real world LAN.

Log into Rancher, select virtualization and the harvester-small cluster, click "Networks" then "VM Networks" and Create.

I named my network "untagged" and left it in the "default" namespace for general use.

Its a type "UntaggedNetwork" (hence the name "untagged") and its on the "Cluster Network" "mgmt" which is the ethernet port of my Harvester nodes.

Cloud-init

https://docs.harvesterhci.io/v1.1/vm/create-vm#cloud-init

https://cloudinit.readthedocs.io/en/latest/

Note OpenStack does drive based and DHCP based cloud-init but Harvester AFAIK only does drive based cloud-init.

One cool feature of Harvester's cloud-init is the SSH keys list seems to support injecting multiple keys into .ssh/authorized_keys.  "Back in the Old Days" using OpenStack, cloud-init only supported one key, this is kind of cool that you can inject multiple keys.

So, lose some features, gain some features.

When creating a VM, the cloud-init config can be manually modified in "Advanced Options" "Cloud Config".

In theory it would be possible to set the serial console password for the ubuntu user or set the VM hostname, but in practice it was not possible.

The link for network config options in the Rancher UI is dead.  

https://cloudinit.readthedocs.io/en/latest/topics/network-config-format-v1.html

The proper link seems to be

https://cloudinit.readthedocs.io/en/latest/reference/network-config-format-v1.html

I submitted a bug report to Harvester at

https://github.com/harvester/harvester/issues/3528

Note this bug was fixed and closed out on March 1st so if you update your Harvester after that date its probably good now.  I write these blog posts some weeks in advance...

If you provide no Network Data, you get a DHCP address, which is acceptable for testing but not useful for production.  Let's design a Network Data for a statically assigned IP address.

version: 1
config:
  - type: physical
    name: enp1s0
    subnets:
      - type: static
        address: 10.10.202.202/16
        gateway: 10.10.1.1
        dns_nameservers:
          - 10.10.7.3
          - 10.10.7.4
        dns_search:
          - cedar.mulhollon.com

cloud-init is an incredibly expensive piece of software to use, because the only feedback you'll get upon any errors is a boot log message "Invalid cloud-config provided: Please run 'sudo cloud-init schema --system' to see the schema errors."  Of course its impossible to run that because if cloud-init doesn't work all user access is blocked, no SSH keys are set and no passwords are set for the serial console.  So best of luck finding your typo or indentation error LOL.

As per the above, I was unable to use cloud-init to set the 'ubuntu' user password for the serial console, or set the VM hostname.  Perhaps there is an incompatibility in the password hash CLI program, or the requirements have changed in an undocumented manner, although I attempted to use the plain_text_password option which also failed.  The online docs that are guaranteed to work with Ubuntu 14 (support ended in 2019) have a disclaimer that those instructions are known not to work in newer installs.  Various attempts to set the hostname silently failed, it was impossible to get a working output from "hostname -f".  I tried using both the "helper" functions in cloud-init and "manually" running various command line commands using runcmd and nothing worked.  Sometimes automation that saves seconds takes too many hours to set up, especially if it has an awful UI.  So just log in as "ubuntu" over SSH and set a password for the serial console manually, and set the hostname manually.  Unfortunate, but necessary.

Create a VM

https://docs.harvesterhci.io/v1.1/vm/create-vm

Log into Rancher, select virtualization and the harvester-small cluster, click "Virtual Machines" then "Create".

The selected namespace will be "default".  I will change the NS to "vm" which is part of the "experimental" project.  How does one know "vm" is part of the "experimental" project in the Create VM UI?  That's an excellent question.  I intend to use a naming strategy for production deployment that looks like if the project name is "projectname" then the NS name will be "projectname-NS" not "NS".  Part of the design advantage of having projects was to permit the re-used of NS names across multiple projects so this naming strategy negates one of the original design goals, but if its unusable, then whatever, do what works as a pragmatism strategy.

I named the test VM "test".

In the "Basics" tab I will provide 1 CPU, 4 GB ram, and the "vince" SSH key.  As per above, "real" production would use the "ansible" key for automated provisioning but using the "vince" key makes it easy to simply log in as "ubuntu@someaddrs" from my "vince" account.

In the "Volumes" tab I select the Ubuntu image as my 10 GB disk image.  I see in the "Type" dropdown there's an option to add a cdrom, so I could install from cdrom if I don't have a ready to use cloud image.  Fun as a virtualized bare metal install would be, which I've done many a time on VMware and OpenStack, in today's experiment, we'll use Ubuntu's cloud image.

In the "Networks" tab the default is type "masquerade" on the management Network.  I will change the type to bridge as I will eventually be using this to host DHCP servers, among other things.  Also I have not experimented with filtering (if any) on a masquerade type network.  The "Network" has to be changed from the internal "management Network", to "default/untagged".

In "Node Scheduling" I do not intend to lock down to any specific node.  Its interesting to look at the rule engine for scheduling.  I could set a key to force certain workloads to certain nodes.   There does not seem to be a facility like "affinity" or "anti-affinity" rules like in VMware, which is too bad.

In "Advanced Options" I see the OS Type was autodetected as "Ubuntu", cool.  See the above cloud-init section to cut and paste in the User Data and Network Data.

Click "Create" and wait.  The UI stopped for a minute or two but stabilized rapidly...

First Five minutes with a new VM

The "test" VM is in status "running" on node "harvester-small-2".

In Rancher the operational tasks for a VM are in the "three dots" menu.  Start, stop, reboot, snapshot, migrate, etc.

First thing I looked at was the logs.  Note these are "Harvester" logs not VM logs.  Lots of lines to research later on.  Main thing I notice is every five seconds I see this log message:

"{"component":"virt-launcher","level":"warning","msg":"Domain id=1 name='vm_test' uuid=feec480d-31a1-59fa-9199-a330c83aa404 is tainted: custom-ga-command","pos":"qemuDomainObjTaintMsg:6382","subcomponent":"libvirt","thread":"30","timestamp":"2023-02-22T20:11:36.552000Z"}"

The timestamp cannot be cut and pasted from a log message, which is annoying.

There is a "console" drop down for webvnc or serial console emulation.  Both seem to work well with this Ubuntu image.  Note that the "ubuntu" does not have a password set, have to configure that after logging in via SSH.

Troubleshooting Lore

Here's some troubleshooting lore, some of which might even be true.

Its possible to wedge into a situation where you can't log in as a static IP address and can't log in as serial because cloud init isn't working.  Seems frustrating.  The solution was to boot up as DHCP, sudo passwd ubuntu, then verify serial console is working, THEN mess around with cloud init network config while logging in over the serial console trying to get static IP addresses working.

I believe network data and/or the user data might only be read on first boot unless something like sudo cloud-init clean is run.

Wednesday, March 8, 2023

Rancher Suite K8S Adventure - Chapter 018 - Tour Harvester Cluster inside Rancher

Rancher Suite K8S Adventure - Chapter 018 - Tour Harvester Cluster inside Rancher

A travelogue of converting from OpenStack to Suse's Rancher Suite for K8S including RKE2, Harvester, kubectl, helm.

This is similar to Chapter 016 where we toured the Harvester UI directly.  Today we compare the differences between a Harvester cluster as seen in Rancher vs the direct Harvester UI.

Log in to Rancher and select "Virtualization Management" from the hamburger menu.  "harvester-small" is the only HCI cluster at this time, click it.

Speed

The first thing I notice overall is the Rancher web UI is dramatically faster.

Multiuser RBAC Auth

Note that the "direct" web UI for Harvester has exactly one user, admin, with superuser privs. The Rancher interface could have multiple users, perhaps fifty, accessing multiple clusters, perhaps ten, with extensive RBAC options for each cluster. In rancher I use admin only to set things up, then I add a user for myself "vince". The Rancher UI for the cluster has an addition left side menu option "Cluster Members" and as an admin user I click "Add" and add myself as a cluster owner. The option for "Custom" permissions provides fine grained roles for cluster access.

Namespaces and Projects

The Rancher "Projects/Namespaces" menu corresponds to the Harvester "Namespaces" menu.  Note that Harvester has no direct concept of Rancher Projects, obviously.  Rancher automatically comes with two projects, "Default" which is prepopulated with Harvester's "Default" namespace, and "Not in a Project" (well, "not in a project" is not really a project, but whatever) and that project is prepopulated with the "harvester-public" namespace.  Note that you can click-thru a namespace in Rancher, such as "harvester-public" and see its resources, it has configmaps and secrets and vm templates and stuff like that.  However in the Harvester web UI you can not click thru and look at the stuff in a namespace.  Probably the weirdest difference I can find between Namespace UI elements is Rancher does not display a "Download YAML" button for a namespace until you checkmark at least one namespace, whereas the Harvester UI displays a grayed out "Download YAML" button until a checkbox for a NS is clicked.  So don't panic if you can't find the YAML download in Rancher, just remember to select a NS first before the button will appear...

Versions

Probably the funniest minor difference is the lower right corner of the screen reports the Rancher version on Rancher and the Harvester version on Harvester.  Conceptually I initially expected the Rancher screen to display the Harvester version when I clicked thru into the Harvester cluster.

Aside from the above differences, the UIs are more or less identical and going forward I will always use the Rancher web UI to control Harvester, although I'll keep Harvester in mind for emergency type access, perhaps if Rancher crashes or something like that.