Skip to content

Helm Install

Complete a Helm install of CubeSandbox on an existing Kubernetes cluster.

Install process

text
① Cluster ready

② Label nodes (and role taints)

③ Prepare values

④ helm upgrade --install

⑤ Verify

1. Prerequisites

ItemRequirement
Kubernetes≥ v1.24+
Toolskubectl, Helm v3.10+
StorageA usable StorageClass in the cluster
Image pullNodes can reach the Internet to pull images

Node roles

CubeSandbox uses labels to distinguish two kinds of nodes:

RoleSuggested sizePurpose
Control-plane node≥ 1; production suggests 3+; ≥ 4C8GRuns Cube's control-plane components
Compute node≥ 1; suggest 16C32G+Runs sandboxes

Recommendation

Use dedicated machines for compute nodes.

What is PVM? (Technical background)

PVM (Pagetable-based Virtual Machine) is a pagetable-based nested virtualization framework built on top of KVM. Unlike traditional nested virtualization, PVM does not rely on the host hypervisor exposing Intel VT-x / AMD-V hardware virtualization extensions to the guest. Instead, it performs privilege-level switching and memory virtualization inside the guest kernel layer through shared-memory regions and shadow page tables, and is fully transparent to the host hypervisor.

PVM was originally proposed in the paper PVM: Efficient Shadow Paging for Deploying Secure Containers in Cloud-native Environment. Tencent Cloud has since done substantial feature and performance work on top of it, fixed many bugs, and open-sourced the upstreamed kernel changes in OpenCloudOS Kernel for the community. Over the years we have deployed a large fleet of PVM instances in Tencent Cloud production, and its reliability has been validated in production.


2. Confirm the cluster is ready

bash
kubectl get nodes -o wide
kubectl get storageclass # Must not be empty
helm version --short # ≥ v3.10

Note the node names you will use later. Confirm at least one node for the control plane and another for the compute plane.

3. Label nodes and roles

Under normal circumstances, add at least 2 machines to Kubernetes and deploy with the control plane and compute plane separated:

Run the following commands to label the relevant nodes:

bash
# Control-plane nodes
kubectl label nodes <control-node> cube.tencent.com/cube-control=true --overwrite

# Compute nodes (run sandboxes)
kubectl label nodes <compute-node> cube.tencent.com/cube-node=true --overwrite

# If the compute node is not a bare-metal server, you also need to apply this label, and PVM will be installed automatically
kubectl label nodes <pvm-compute-node> cube.tencent.com/allow-pvm-bootstrap=true --overwrite

Set role taints (to keep unrelated workloads off these nodes):

bash
# Optional
# kubectl taint nodes <control-node> cube.tencent.com/control=true:NoSchedule --overwrite

# Required for compute nodes
kubectl taint nodes <compute-node> cube.tencent.com/compute=true:NoSchedule --overwrite

Hint

For a single-node setup, we recommend deploying via Quick Start instead of Kubernetes.

Single-node labels and taints

If your Kubernetes cluster has only one machine, you can try the following labels and taints (not recommended):

bash
export NODE=<your-node-name>

kubectl label nodes "$NODE" \
  cube.tencent.com/cube-control=true \
  cube.tencent.com/cube-node=true \
  cube.tencent.com/allow-pvm-bootstrap=true \
  --overwrite

kubectl taint nodes "$NODE" cube.tencent.com/control=true:NoSchedule --overwrite
kubectl taint nodes "$NODE" cube.tencent.com/compute=true:NoSchedule --overwrite

4. Prepare the values file

Create the configuration file:

bash
cp deploy/kubernetes/chart/runtime-values.example.yaml runtime-values.yaml

Edit it to fit your needs. At minimum, fill in the following:

yaml
cubeProxy:
  advertiseIP: "10.0.1.10"
  domain: "cube.app"
  tls:
    mode: selfSigned         # Trial; production: existingSecret / certManager

# Required
mysql:
  host: ""                   # Empty = use built-in MySQL
  password: "replace-me-mysql-password"
  rootPassword: "replace-me-mysql-root-password"
redis:
  host: ""
  password: "replace-me-redis-password"

More configuration

Control-plane PVC configuration

See 8.1 Control-plane PVC configuration.

Compute node data disk

Cube writes data to the host at /data/cubelet, which must be an XFS filesystem. For a smoother first-time experience, we create a 25 GB loopback image file and mount it by default.

For production deployments, or to adjust related configuration, see 8.2 Compute node data disk configuration.

Compute node networking

Decide before you deploy

Once sandboxes are running, the cube-node Pod on a compute node must not be recreated: recreation destroys the network namespace that sandbox network devices live in, and all sandbox networking on that node breaks — inbound and outbound. For production, deploy cube-node with hostNetwork so Pod recreation no longer changes the netns. Details and caveats (including NetworkPolicy): 8.3 cube-node networking and Pod recreation.

5. Helm install

Users in Mainland China

Add -f deploy/kubernetes/chart/values-cn.yaml to whichever install command you use below to pull images from the China registry.

Run the following commands from the repository root to complete the install:

Because this process downloads many resources and runs initialization actions, it can take some time depending on your machine performance and network conditions.

Standard multi-node deployment:

bash
# Standard multi-node
helm upgrade --install cube ./deploy/kubernetes/chart \
  -n cube-system \
  --create-namespace \
  -f runtime-values.yaml \
  --wait \
  --timeout 90m

Tencent Cloud TKE:

bash
helm upgrade --install cube ./deploy/kubernetes/chart \
  -n cube-system \
  --create-namespace \
  -f deploy/kubernetes/chart/values-tke.yaml \
  -f runtime-values.yaml \
  --wait \
  --timeout 90m

Single-node trial:

bash
helm upgrade --install cube ./deploy/kubernetes/chart \
  -n cube-system \
  --create-namespace \
  -f deploy/kubernetes/chart/values-single-node.yaml \
  -f runtime-values.yaml \
  --wait \
  --timeout 90m

6. Verify the deployment

bash
# 1) Are Pods Ready?
kubectl get pods -n cube-system -o wide

# 2) Have compute nodes registered with CubeOps?
kubectl exec -n cube-system deploy/cube-cubemastercli -- \
  sh -lc 'cubeopscli --address "$CUBEOPSCLI_ADDRESS" --port "$CUBEOPSCLI_PORT" node list'

# 3) Built-in end-to-end tests (a few minutes)
helm test cube -n cube-system --timeout 20m --logs

Expect:

  • cube-node Ready count = number of nodes labeled cube-node=true
  • All services in the cube-system namespace are Ready
  • helm test passes

Access the Cube WebUI:

bash
kubectl -n cube-system port-forward svc/cube-webui 12088:12088
# Open http://127.0.0.1:12088 in your browser

7. Uninstall

bash
helm uninstall cube -n cube-system
kubectl delete namespace cube-system

Uninstall does not automatically clean up:

  • Cube labels / role taints on nodes
  • Compute-node hostPath data
  • PVM host kernel changes

8. Advanced configuration

8.1 Control-plane PVC configuration

  • Default: CubeMaster / MySQL / Redis / MinIO use the cluster default StorageClass
  • To specify an SC: set persistence.storageClassName: <name> in runtime-values.yaml
  • Single-node / no CSI: switch to hostPath (see comments in runtime-values.example.yaml)

8.2 Compute node data disk configuration

For production deployments, we recommend provisioning a dedicated data disk for each compute node, formatting it as XFS, and mounting it under /data/cubelet. This gives the best performance and stability. Then configure runtime-values.yaml:

yaml
bootstrap:
  nodeInit:
    dataCubelet:
      loopback:
        enabled: false

If your nodes have no dedicated data disk, by default we mount a 25 GB loop device at /data/cubelet during compute-node initialization.

Only affects the "first" creation

By default, the disk image file is created automatically only when /data/cubelet-xfs.img does not exist. Subsequent config changes do not auto-expand. To resize, manual steps in a maintenance window are required.

To customize the data disk configuration before install:

values pathDefaultDescription
bootstrap.nodeInit.dataCubelet.loopback.enabledtrueWhen true, bootstrap creates and mounts a loopback XFS on the node
bootstrap.nodeInit.dataCubelet.loopback.imagePath/data/cubelet-xfs.imgPath to the image file
bootstrap.nodeInit.dataCubelet.loopback.size25GSize when the image is first created (truncate -s)
yaml
bootstrap:
  nodeInit:
    dataCubelet:
      loopback:
        enabled: true
        size: 200G   # Size to capacity plan; must be less than free space on the filesystem holding the image

8.3 cube-node networking and Pod recreation

In one sentence

While sandboxes are running, the cube-node Pod must not be recreated — recreation breaks network connectivity for all sandboxes on the node (inbound and outbound) with no self-healing. For production, enable hostNetwork at deploy time to avoid this; the trade-off is that NetworkPolicy needs extra handling (see below).

Why the cube-node Pod must not be recreated

Sandbox network devices (TAP devices) and the cubevs hooks live in the network namespace of the cube-node Pod; recreating the Pod destroys that netns:

ItemDescription
TriggerAnything that recreates the Pod: DaemonSet template change, image bump, manual kubectl delete pod, …
ImpactAll sandboxes on the node lose network connectivity — both inbound and outbound
Self-healingNone. The only recovery is to destroy and recreate the affected sandboxes

Compute-plane upgrades are affected for the same reason; see Upgrade.

Recommendation: enable hostNetwork at deploy time

Recommended

Run cube-node with hostNetwork: true before any sandbox is created: the Pod then shares the host netns, which does not change across Pod recreation, so sandbox network devices survive a cube-node rebuild.

What to check when enabling it:

ItemDescription
How to enableThe Chart has no values toggle (security.hostNetwork is rejected by validation); patch the DaemonSet via a Helm post-renderer / Kustomize / a fork of the Chart
DNSAlso set dnsPolicy: ClusterFirstWithHostNet so in-cluster DNS keeps resolving
Port conflictsEnsure cubelet's ports (9998 / 9999 / 9966) do not conflict with other services on the host
Monitoring / firewallsRe-point anything keyed on the Pod IP / Pod CIDR at the node IP instead

Trade-off: NetworkPolicy no longer applies

With hostNetwork, cube-node loses its CNI-assigned Pod identity, so Kubernetes NetworkPolicy (e.g. "can sandboxes access this Service") cannot govern sandbox traffic directly.

Reference implementation: PR #1189 (not yet merged; for reference only)

If you need NetworkPolicy over sandbox traffic to in-cluster Services / Pods, see PR #1189:

  • Only traffic destined for the cluster CIDRs is forwarded through a node-local EgressProxy Pod and SNAT'd to the Proxy Pod IP, so it is governed by your own NetworkPolicies; all other traffic keeps the normal route.
  • The PR also makes hostNetwork the default and implements the full in-place cube-node replacement design.

FAQ quick reference

SymptomWhat to do
Control-plane Pods stay PendingBy default nodes need cube-control=true (or your custom nodeSelector); also check taints / tolerations
cube-node DESIRED=0Are compute nodes labeled cube-node=true?
Helm reports CHANGE_ME_* / validation failureCheck that passwords and placement in runtime-values.yaml are complete
First install is slow / nodes rebootExpected when PVM is enabled; wait for fingerprints to be ready and gates cleared, then check bootstrap / Big Pod

More detail in the other docs in this directory:


Next steps