This is all from a medium size company we don’t use Kubernetes for the scaling, we use it because we can.
It is a simple and stable life, but
ingress-nginx was deprecated, so we have to replace it, probably with traefik
kubenet networking on Azure was deprecated, Azure wants us to re-do our production cluster by 2028
Single Node Cluster works just fine
For years we run single node cluster for the data-warehouse. There were no problems.
Default StorageClass with ZRS and Retain
The default StorageClass may use LRS and reclaimPolicy: Delete. Which means it creates the disks in a specific zone, and it will remove the disks automatically.
You want to create a StorageClass that uses ZRS and reclaimPolicy: Retain.
We used LRS, one day the cluster moved to another zone, a zone different from the disks.
Avoid CPU limits
CPU Limits are tricky and hard to get right. We don’t set them. [more]
One services used up all its CPU time and got stuck in GC pauses.
AKS Reader Role with Metrics
Per default there is no Azure AKS role for developers. We create a new reader role that can access the metrics, pods, configmaps, but not secrets.
{
"permissions": [
{
"actions": [
"Microsoft.Authorization/*/read",
"Microsoft.ContainerService/managedClusters/listClusterUserCredential/action",
"Microsoft.Insights/alertRules/*",
"Microsoft.Resources/subscriptions/operationresults/read",
"Microsoft.Resources/subscriptions/read",
"Microsoft.Resources/subscriptions/resourceGroups/read",
"Microsoft.Support/*"
],
"notActions": [],
"dataActions": [
"Microsoft.ContainerService/managedClusters/apps/controllerrevisions/read",
"Microsoft.ContainerService/managedClusters/apps/daemonsets/read",
"Microsoft.ContainerService/managedClusters/apps/deployments/read",
"Microsoft.ContainerService/managedClusters/apps/replicasets/read",
"Microsoft.ContainerService/managedClusters/apps/statefulsets/read",
"Microsoft.ContainerService/managedClusters/autoscaling/horizontalpodautoscalers/read",
"Microsoft.ContainerService/managedClusters/batch/cronjobs/read",
"Microsoft.ContainerService/managedClusters/batch/jobs/read",
"Microsoft.ContainerService/managedClusters/configmaps/read",
"Microsoft.ContainerService/managedClusters/endpoints/read",
"Microsoft.ContainerService/managedClusters/events.k8s.io/events/read",
"Microsoft.ContainerService/managedClusters/events/read",
"Microsoft.ContainerService/managedClusters/extensions/daemonsets/read",
"Microsoft.ContainerService/managedClusters/extensions/deployments/read",
"Microsoft.ContainerService/managedClusters/extensions/ingresses/read",
"Microsoft.ContainerService/managedClusters/extensions/networkpolicies/read",
"Microsoft.ContainerService/managedClusters/extensions/replicasets/read",
"Microsoft.ContainerService/managedClusters/limitranges/read",
"Microsoft.ContainerService/managedClusters/namespaces/read",
"Microsoft.ContainerService/managedClusters/networking.k8s.io/ingresses/read",
"Microsoft.ContainerService/managedClusters/networking.k8s.io/networkpolicies/read",
"Microsoft.ContainerService/managedClusters/persistentvolumeclaims/read",
"Microsoft.ContainerService/managedClusters/pods/read",
"Microsoft.ContainerService/managedClusters/policy/poddisruptionbudgets/read",
"Microsoft.ContainerService/managedClusters/replicationcontrollers/read",
"Microsoft.ContainerService/managedClusters/resourcequotas/read",
"Microsoft.ContainerService/managedClusters/serviceaccounts/read",
"Microsoft.ContainerService/managedClusters/services/read",
"Microsoft.ContainerService/managedClusters/metrics/read",
"Microsoft.ContainerService/managedClusters/metrics.k8s.io/nodes/read",
"Microsoft.ContainerService/managedClusters/metrics.k8s.io/pods/read",
"Microsoft.ContainerService/managedClusters/resetMetrics/read",
"Microsoft.ContainerService/managedClusters/apis/metrics.k8s.io/read",
"Microsoft.ContainerService/managedClusters/nodes/read"
],
"notDataActions": []
}
]
}
}Weird Cost Optimisations
on AKS the amount of memory that is available for normal pods is tied to the max-pods configuration, per default it is 110, if you tune it down you get more memory
Azure AKS system pods request a lot of CPU resources (I think it is 1.5 CPU). This can become costly with a lot of small system nodes in your test environments.
Turn off log collection if not needed link
Other Weird Things
You still need a preStop sleep 10 on every deployment, to give the ingress and traffic time to drain.
lifecycle:
preStop:
exec:
command:
- /bin/sleep
- '10'We name our Kubernetes ready endpoints /kubernetes/ready to send a clear message, this endpoint belongs to Kubernetes, you respond 200 when you want to receive traffic. The endpoint /health/ready was misunderstood and used for all kind of health checks (thanks Microsoft).
Liveness probes are only used for applications that have a history of hanging/freeze.
readinessProbe:
initialDelaySeconds: 5
periodSeconds: 10
failureThreshold: 3
successThreshold: 1
httpGet:
path: /kubernetes/ready
port: 8080If a deployment has dependencies to secrets they are noted in a label, so the relationship can be queried withe the kubectl cli in scripts.
Nodepool Image Hydration
When we add a new nodepool we download all the images before we move workload over. The bash script
#!/usr/bin/env bash
set -euo pipe fail
# Note: each image must contain the true binary, otherwise it will fail
NAMESPACE="${1:-default}"
echo "Fetching images from namespace: $NAMESPACE"
init_containers=$(kubectl get pods -n "$NAMESPACE" \
-o jsonpath="{.items[*].spec['initContainers','containers'][*].image}" \
| tr ' ' '\n' | sort -u \
| jq -Rn '[inputs | select(length > 0)] | to_entries | map({
name: ("prepuller-" + (.key + 1 | tostring)),
image: .value,
command: ["true"]
})')
echo "Found $(echo "$init_containers" | jq length) images, generating prepuller.json"
jq -n --argjson ic "$init_containers" '{
apiVersion: "apps/v1",
kind: "DaemonSet",
metadata: { name: "prepuller" },
spec: {
selector: { matchLabels: { name: "prepuller" } },
template: {
metadata: { labels: { name: "prepuller" } },
spec: {
initContainers: $ic,
containers: [{
name: "pause",
image: "registry.k8s.io/pause:latest"
}]
}
}
}
}' > prepuller.json
echo "Applying the generated prepuller.json to start pulling images..."
echo " kubectl apply -n $NAMESPACE -f prepuller.json"
kubectl apply -n $NAMESPACE -f prepuller.json
echo "Waiting for the prepuller DaemonSet to be ready..."
echo " kubectl wait -n $NAMESPACE --for=condition=ready pod -l name=prepuller --timeout=300s"
kubectl wait -n $NAMESPACE --for=condition=ready pod -l name=prepuller --timeout=300s
echo "Deleting puller DaemonSet..."
echo " kubectl delete -n $NAMESPACE daemonset prepuller"
kubectl delete -n $NAMESPACE daemonset prepuller
rm prepuller.json
echo "Finished."Simple Helm Deployments
We keep our Kubernetes definition simple, only pass in the image URL during deployment. Helm does a good job and with the atomic option it rolls back on any error during deployment.
In bigger projects we generate the Helm Charts (which are simple Kubernetes YAML files) from custom YAML templates, with configurations, written in YAML.
The ingress configurations (the configuration of what URL maps to which deployment) used to live next to the code, but we decouple it to prevent older branches from re-deploying older ingress.
We are a small company so simple solution go a long way.
Helm Chart
An example of a very simple and flat helm chart.
values.yaml
deploymentName: NAME
TargetEnvironment: prod
port: 4200
replicas: 2
requests:
cpu: 500m
memory: 2000Mi
limits:
memory: 2000Mi
deployment.txt (it is not valid yaml)
apiVersion: apps/v1
kind: Deployment
metadata:
name: '{{ .Values.deploymentName }}'
labels:
app.kubernetes.io/name: '{{ .Values.deploymentName }}'
app.kubernetes.io/part-of: backend
spec:
replicas: {{ .Values.replicas }}
revisionHistoryLimit: 3
selector:
matchLabels:
app.kubernetes.io/name: '{{ .Values.deploymentName }}'
template:
metadata:
labels:
app.kubernetes.io/name: '{{ .Values.deploymentName }}'
app.kubernetes.io/part-of: backend
spec:
automountServiceAccountToken: false
enableServiceLinks: false
terminationGracePeriodSeconds: 60
securityContext:
runAsNonRoot: true
containers:
- image: '{{ .Values.image }}'
imagePullPolicy: IfNotPresent
name: '{{ .Values.deploymentName }}'
env:
- name: TARGET_ENVIRONMENT
value: '{{ .Values.TargetEnvironment }}'
- name: PORT
value: '{{ .Values.port }}'
- name: POD_DEPLOYMENT_NAME
valueFrom:
fieldRef:
fieldPath: metadata.labels['app']
- name: POD_NAME
valueFrom:
fieldRef:
fieldPath: metadata.name
- name: POD_NAMESPACE
valueFrom:
fieldRef:
fieldPath: metadata.namespace
- name: POD_MEMORY_LIMIT
valueFrom:
resourceFieldRef:
resource: limits.memory
divisor: 1Mi
ports:
- containerPort: {{ .Values.port }}
lifecycle:
preStop:
exec:
command:
- /bin/sleep
- '10'
readinessProbe:
initialDelaySeconds: 1
periodSeconds: 5
failureThreshold: 3
successThreshold: 1
httpGet:
path: /kubernetes/ready
port: {{ .Values.port }}
resources:
requests:
cpu: {{ .Values.requests.cpu }}
memory: {{ .Values.requests.memory }}
limits:
memory: {{ .Values.limits.memory }}
service.txt
apiVersion: v1
kind: Service
metadata:
name: '{{ .Values.deploymentName }}'
labels:
app.kubernetes.io/part-of: backend
spec:
type: ClusterIP
ports:
- port: 80
targetPort: {{ .Values.port }}
selector:
app.kubernetes.io/name: '{{ .Values.deploymentName }}'
