One Model, Many Tenants: Testing a Shared P2P Cache Across Virtual Kubernetes Clusters
A reproducible lab with Dragonfly and vCluster shows how a platform-owned model distribution layer can reuse a public model payload across tenants without adding node-level privileges to tenant workloads.
Multi-tenant Kubernetes platforms often give each internal team its own API boundary. That separation helps teams move independently, but it does not stop them from requesting the same large AI model files. When every tenant downloads the same weights directly from a model hub, the platform repeats network transfers and stores duplicate data.
A cache inside every tenant preserves ownership, but it also limits reuse. A node-level cache can share data more broadly, yet tenants should not need hostPath mounts, host networking, or privileged containers to use it. I built a small lab to test whether a platform-owned peer-to-peer cache could serve the same model payload across two virtual clusters while keeping those node-level controls outside the tenant workload.
The Platform-Owned Data Path
Dragonfly ran only on the host cluster in manager-less mode. Schedulers and seed peers handled task coordination, while a client DaemonSet ran on each node. Each client exposed an HTTP proxy on port 4001 through host networking. The platform team owned these components and the cache.
Tenant A and Tenant B each used a separate open source vCluster API server. Their model-download Jobs were ordinary non-root Pods. The manifest did not use hostPath, hostNetwork, hostPID, hostIPC, privileged mode, or added Linux capabilities. The Job read its node IP through the Kubernetes downward API and sent the model payload through the node-local Dragonfly proxy.
The Job first sent a HEAD request to Hugging Face to resolve the model file’s signed CDN URL. It then downloaded the payload through Dragonfly. The proxy configuration removed selected signed query parameters from the cache identity so both tenants mapped the same content path to the same Dragonfly task. This detail matters because a changing signed URL can otherwise prevent reuse.
Figure 1. Host-managed Dragonfly data plane shared by two vCluster tenants.
What the Run Showed
Tenant A ran on the host control-plane node. It downloaded 267,954,768 bytes and produced SHA-256 checksum 5e3f1108e3cb34ee048634875d8482665b65ac713291a7e32396fb18f6ff0063. Its observed wall-clock interval was about 11 seconds.
I then placed Tenant B on dragonfly-host-worker2 so the second request could not be satisfied by the same node’s local cache. Dragonfly assigned both requests the same task ID. For Tenant B, the scheduler returned the control-plane client and a seed peer as parents. The control-plane client logged that it sent its existing pieces to Worker 2. Tenant B received the same 267,954,768-byte file and produced the same checksum. Its observed interval was about three seconds.
The shorter interval was not the proof. The useful evidence was the shared task ID, the parent list, the explicit remote-peer upload log, the different host nodes, and the matching checksum. Those records establish remote-peer delivery of the model payload. I did not independently meter total origin bytes, so the result should not be described as eliminating every request to Hugging Face.
Platform Engineering Lessons
First, the cache can be treated as a platform service. Tenants use a narrow data path, while the platform owns deployment, capacity, observability, and upgrades. Second, cache identity is part of the architecture. Signed model-hub URLs must be normalized carefully, and the rule used in this lab should not be copied to private or gated models without separate authorization and cache-isolation analysis.
Third, transfer-source logs matter more than timing. A faster second pull may indicate cache reuse, but it does not prove whether data came from a local cache, a remote peer, or the origin. Finally, API separation is not the same as hostile-tenant security. This lab tested ordinary tenant workloads with restricted manifests. It did not evaluate adversarial behavior.
Alternatives and Production Rule
A platform team could also use a per-node registry mirror, a pull-through cache, or a shared persistent-volume cache. The right choice depends on artifact formats, trust boundaries, storage design, and operational ownership. I chose Dragonfly and vCluster here to test cross-tenant P2P reuse while keeping the cache under platform control.
The production rule I would carry forward is simple: centralize repeated artifact distribution as a platform capability, expose only the smallest tenant-facing interface, and verify the actual transfer source. Start with public artifacts and explicit cache keys. Treat private-model authorization, eviction, failure recovery, and capacity planning as separate production requirements.
Author Bio
Pavan Madduri is a Senior Cloud Platform Engineer at W.W. Grainger and the elected Tech Lead of CNCF’s TAG Workloads Foundation (2026 to 2028). He contributes to Dragonfly, a graduated CNCF project, and is a vCluster Ambassador and CNCF Golden Kubestronaut.
Disclosure: the author is a member of the vCluster Ambassador Program. This article uses only open source vCluster features and was not reviewed or sponsored by vCluster Labs.


