Manual snapshots (Part 2) are for the moment before a risky change. Day-to-day protection is a schedule. NDK splits that into three objects, and this part builds them for the journal application, watches the first run fire, and then looks at a plan that has been running on our lab for 25 days to see what retention really does.

Lab: NKP 2.18, NDK 2.3.0, all outputs from 2026-09-19.

All manifests in this part are in github.com/Fen0l/ndk-examples, the same files we applied on the lab.

Three objects, one job each

Object Answers Scope
JobScheduler When? namespace, reusable by several plans
ProtectionPlan What kind of protection, how many to keep, replicate where? namespace
AppProtectionPlan Which Application gets which plans? namespace, one per Application

The split looks heavy for a first plan. It pays off when you have twenty applications on the same "hourly, keep 24" policy: one scheduler, one plan, twenty bindings.

The 60-minute floor

Before writing the schedule, know the limit. NDK refuses anything more frequent than once an hour, on every schedule type. We tried both:

output
The JobScheduler "journal-every-5m" is invalid: spec.cronSchedule: Invalid value: "*/5 * * * *":
Cron schedule of lower interval than 60 minutes is not allowed

The JobScheduler "journal-30m" is invalid: spec.interval.minutes: Invalid value: 30:
Interval time must be a positive integer greater than or equal to 60

That is a validating webhook, so the object never lands in etcd. If your RPO requirement is below one hour, snapshots are the wrong tool and you are looking at NearSync or sync replication, which need a PE-level topology we do not have in the lab.

The manifests

yaml
apiVersion: scheduler.nutanix.com/v1alpha1
kind: JobScheduler
metadata:
  name: journal-hourly
  namespace: ndk-howto
spec:
  interval:
    minutes: 60
  startTime: "2026-09-19T14:05:00Z"
  timeZoneName: Etc/UTC
---
apiVersion: dataservices.nutanix.com/v1alpha1
kind: ProtectionPlan
metadata:
  name: journal-local-hourly
  namespace: ndk-howto
spec:
  protectionType: async
  scheduleName: journal-hourly
  retentionPolicy:
    retentionCount: 2
---
apiVersion: dataservices.nutanix.com/v1alpha1
kind: AppProtectionPlan
metadata:
  name: journal-protection
  namespace: ndk-howto
spec:
  applicationName: journal
  protectionPlanNames:
    - journal-local-hourly

A few notes on the fields:

  • startTime is optional. Without it, an interval schedule starts counting from the moment you create the object. We set it three minutes ahead so we could watch the first run. daily, weekly, monthly and cronSchedule are the other schedule types, all with the same 60-minute floor.
  • protectionType: async with no replicationConfigs means local snapshots only. Adding a replicationConfigs list with a replicationTargetName makes every scheduled snapshot replicate (Part 4).
  • retentionCount is between 1 and 15. It counts successful snapshots on this cluster only.
  • The AppProtectionPlan is where protection actually starts. Nothing happens until it exists.
bash
kubectl apply -f 05-scheduled-protection.yaml
kubectl -n ndk-howto get jobscheduler,protectionplan,appprotectionplan
output
NAME                                                LASTACTIVATION   NEXTACTIVATION
jobscheduler.scheduler.nutanix.com/journal-hourly                    2026-09-19T14:05:00Z

NAME                                                           SCHEDULE-NAME    RETENTION-COUNT   AVAILABLE   DEGRADED   PROTECTION-TYPE
protectionplan.dataservices.nutanix.com/journal-local-hourly   journal-hourly   2                 True        False      async

NAME                                                            APPLICATIONNAME   PROTECTIONPLANS-APPLIED    AVAILABLE   DEGRADED
appprotectionplan.dataservices.nutanix.com/journal-protection   journal           ["journal-local-hourly"]   True        False

The NEXTACTIVATION column is your first check. If it is empty, the schedule spec was not understood.

Binding a plan also changes the Application. Its finalizers went from one to two:

output
["dataservices.nutanix.com/app","dataservices.nutanix.com/app-protection-plan"]

This is what we pointed at in Part 1: an Application with a plan cannot be deleted until the AppProtectionPlan is gone. Delete in reverse order of creation.

The first run

bash
kubectl -n ndk-howto get applicationsnapshot -w

At 14:05:28Z, 28 seconds after the scheduled time:

output
NAME                               AGE   READY-TO-USE   BOUND-SNAPSHOTCONTENT                      SNAPSHOT-AGE   CONSISTENCY-TYPE
journal-before-change              20m   true           asc-fd129a47-fc53-4572-91b3-2c473555a3c9   19m            CrashConsistent
journal-f40c53315adf8d8c-1c72d2d   27s   false          asc-f5ba48b0-929c-411e-898d-9bf893bdf11c

The snapshot's creationTimestamp is 2026-09-19T14:05:00Z. Not 14:05:03, not 14:05:12. The scheduler fires on the second, and we saw the same on the 25-day-old plan below (every run at exactly 16:07:00Z). Half a minute later it was READY-TO-USE: true, CrashConsistent.

The scheduler and the binding both record the run:

bash
kubectl -n ndk-howto get jobscheduler journal-hourly -o jsonpath='{.status}'
kubectl -n ndk-howto get appprotectionplan journal-protection -o jsonpath='{.status.protectionPlanExecutionStatus}'
output
{"lastActivation":"2026-09-19T14:05:00Z","lastUpdatedAt":"2026-09-19T14:05:00Z","nextActivation":"2026-09-19T15:05:00Z"}
[{"lastExecutionTime":"2026-09-19T14:05:00Z","lastScheduledExecutionTime":"2026-09-19T14:05:00Z","protectionPlanName":"journal-local-hourly"}]

How to tell a scheduled snapshot from a manual one

The generated name (journal-f40c53315adf8d8c-1c72d2d) is application name, a hash, and a counter. More useful are the labels NDK puts on it:

bash
kubectl -n ndk-howto get applicationsnapshot journal-f40c53315adf8d8c-1c72d2d -o jsonpath='{.metadata.labels}' | jq .
json
{
  "dataservices.nutanix.com/app-protection-plan": "journal-protection",
  "dataservices.nutanix.com/application-name": "journal",
  "dataservices.nutanix.com/application-namespace": "ndk-howto",
  "dataservices.nutanix.com/protection-plan": "journal-local-hourly"
}

(plus the UIDs of both plans). So kubectl get applicationsnapshot -l dataservices.nutanix.com/protection-plan=journal-local-hourly lists exactly what one plan produced. And spec.expiresAfter is empty on a scheduled snapshot: retention is the plan's job, which is why the webhook in Part 2 only demands expiresAfter on manual ones.

Retention, observed over 25 days

One hourly run does not show retention. Our older ndk-demo application does. It has had a daily plan with retentionCount: 3 since 2026-08-24, and nobody touched it since.

bash
kubectl -n ndk-demo get jobscheduler,protectionplan
output
NAME                                            LASTACTIVATION         NEXTACTIVATION
jobscheduler.scheduler.nutanix.com/demo-daily   2026-09-18T16:07:00Z   2026-09-19T16:07:00Z

NAME                                                      SCHEDULE-NAME   RETENTION-COUNT   AVAILABLE   DEGRADED   PROTECTION-TYPE
protectionplan.dataservices.nutanix.com/demo-local-plan   demo-daily      3                 True        False      async

25 daily runs. Here is what is left:

bash
kubectl -n ndk-demo get applicationsnapshot -o custom-columns='NAME:.metadata.name,CREATED:.metadata.creationTimestamp,READY:.status.readyToUse,CONS:.status.consistencyType'
output
NAME                                CREATED                READY   CONS
ndk-demo-e51de6e006b519b4-1c6c2c7   2026-08-31T16:07:00Z   false   <none>
ndk-demo-e51de6e006b519b4-1c6d947   2026-09-04T16:07:00Z   false   <none>
ndk-demo-e51de6e006b519b4-1c71cc7   2026-09-16T16:07:00Z   true    CrashConsistent
ndk-demo-e51de6e006b519b4-1c72267   2026-09-17T16:07:00Z   true    CrashConsistent
ndk-demo-e51de6e006b519b4-1c72807   2026-09-18T16:07:00Z   true    CrashConsistent

Retention works: exactly three ready snapshots, the three most recent. The other 20 runs are gone, pruned as newer ones arrived.

But there are five objects, not three. Two runs, on 08-31 and 09-04, never reached READY. Their content objects say why:

bash
kubectl get applicationsnapshotcontent asc-fc169e04-bc6d-4d4c-bf30-23f8a8e16614 -o jsonpath='{.status.conditions}' | jq -c '.[] | {type,status,reason}'
output
{"type":"Progressing","status":"False","reason":"VolumeSnapshotCreationFailedDueToBlockVolumes"}
{"type":"AppConfigAcquired","status":"True","reason":"AcquiredAppConfig"}
{"type":"VolumeSnapshotsCreated","status":"False","reason":"VolumeSnapshotCreationFailedDueToBlockVolumes"}

The full message carries the Prism API response: AUTHENTICATION_REQUIRED, a 401. On those two evenings, Prism Central rejected the service account for a few minutes (we saw the same thing during this write-up, it cleared on its own). NDK asked for the volume snapshot, got a 401, and marked the run failed. No retry.

Two conclusions from that, and they matter more than the manifests:

  1. Failed snapshots are not retried and not pruned. They sit outside the retention count, with readyToUse: false, until you delete them. Over a year, a plan with occasional failures accumulates objects. readyToUse is a status field, so a label or field selector cannot find them; list them with jq and delete by name:
bash
kubectl -n ndk-demo get applicationsnapshot -o json \
  | jq -r '.items[] | select(.status.readyToUse != true) | .metadata.name'
output
ndk-demo-e51de6e006b519b4-1c6c2c7
ndk-demo-e51de6e006b519b4-1c6d947
bash
kubectl -n ndk-demo delete applicationsnapshot ndk-demo-e51de6e006b519b4-1c6c2c7 ndk-demo-e51de6e006b519b4-1c6d947

Check the list before you pipe it into a delete. A snapshot that is still in progress also has readyToUse: false. 2. The plan reported healthy the whole time. AVAILABLE: True, DEGRADED: False, on both the ProtectionPlan and the AppProtectionPlan, before, during and after the failures. Those conditions describe the plan's configuration, not its results. If your monitoring watches plan status, it will never page.

What to watch instead: the age of the newest applicationsnapshot with readyToUse: true per application. If it is older than your schedule interval plus a margin, protection is broken. Part 5 looks at what NDK exposes to Prometheus today, and it is less than you would hope.

Changing a plan: three webhooks you will meet

We wanted to add replication to the running plan. That turned into a tour of NDK's admission rules, all observed on the lab.

A ProtectionPlan is immutable.

bash
kubectl -n ndk-howto patch protectionplan journal-local-hourly --type=merge \
  -p '{"spec":{"replicationConfigs":[{"replicationTargetName":"demo-wkl-02"}]}}'
output
The ProtectionPlan "journal-local-hourly" is invalid: spec: Invalid value: Spec is immutable for protectionPlan.dataservices.nutanix.com

Retention, schedule, replication: none of it can change after creation. You create a new plan.

Plans can be added to an AppProtectionPlan, not removed.

bash
kubectl -n ndk-howto patch appprotectionplan journal-protection --type=merge \
  -p '{"spec":{"protectionPlanNames":["journal-hourly-to-demo-wkl-02"]}}'
output
admission webhook "vappprotectionplan.kb.io" denied the request: removing protection plans from an AppProtectionPlan is not allowed: [journal-local-hourly]. Delete the AppProtectionPlan instead to remove protection

A bound ProtectionPlan will not delete.

bash
kubectl -n ndk-howto delete protectionplan journal-local-hourly --timeout=30s
output
protectionplan.dataservices.nutanix.com "journal-local-hourly" deleted from ndk-howto namespace
error: timed out waiting for the condition on protectionplans/journal-local-hourly

The object sits in Terminating with the finalizer dataservices.nutanix.com/app-protection-plan-journal-protection, named after the binding that holds it. It goes away the moment the binding does.

So the working sequence to replace a plan is:

bash
kubectl apply -f 09-replicated-plan.yaml                      # new ProtectionPlan
kubectl -n ndk-howto delete appprotectionplan journal-protection   # old binding (old plan finalizes now)
kubectl apply -f 10-appprotectionplan-replicated.yaml         # new binding

Five seconds after the binding was deleted, the old plan was gone and the Application was back to a single finalizer. Existing snapshots stayed: they belong to the Application, not to the plan. Then one more thing happened that is worth knowing.

Binding a plan can snapshot immediately

Our first binding, at 14:02 with a startTime three minutes ahead, did nothing until 14:05:00. The second binding, at 14:12, against a scheduler whose startTime was already in the past, took a snapshot right away:

output
NAME                               CREATED                READY   PLAN
journal-316d190781390250-1c72d2d   2026-09-19T14:12:53Z   false   journal-hourly-to-demo-wkl-02
journal-f40c53315adf8d8c-1c72d2d   2026-09-19T14:05:00Z   true    journal-local-hourly
journal-before-change              2026-09-19T13:45:13Z   true    <none>

Reasonable behaviour (a newly protected application should not wait an hour for its first copy), but plan for it on a large application, and do not rebind twenty applications at once on a busy afternoon.

Three hours later, the same list showed retention at work on this plan too:

output
NAME                               CREATED                READY   PLAN
journal-316d190781390250-1c72d69   2026-09-19T15:05:00Z   true    journal-hourly-to-demo-wkl-02
journal-316d190781390250-1c72da5   2026-09-19T16:05:00Z   true    journal-hourly-to-demo-wkl-02
journal-f40c53315adf8d8c-1c72d2d   2026-09-19T14:05:00Z   true    journal-local-hourly
journal-before-change              2026-09-19T13:45:13Z   true    <none>

retentionCount: 2, three runs (14:12, 15:05, 16:05), two left: the 14:12 snapshot was pruned when the 16:05 one became ready. The 14:05 snapshot from the deleted plan is still there. Retention only counts snapshots of the plan that made them, so a plan you delete leaves its snapshots behind until they expire or you remove them.

Removing scheduled protection

Same order: binding, plan, scheduler.

bash
kubectl -n ndk-howto delete appprotectionplan journal-protection
kubectl -n ndk-howto delete protectionplan journal-hourly-to-demo-wkl-02
kubectl -n ndk-howto delete jobscheduler journal-hourly

Summary

Scheduled protection in NDK is three objects: a JobScheduler (never more often than hourly), a ProtectionPlan (type, retention, replication, immutable once created), and an AppProtectionPlan that binds a plan to an Application. The schedule fires on the second and retention keeps exactly the count you asked for. What it does not do is retry or clean up failed runs, and the plan's own status stays green through them, so monitor snapshot freshness, not plan conditions.

In Part 4 the same snapshot goes to a second cluster: Remote, ReplicationTarget, and a restore on the other side.