DPSGDStep¶
The differentially private SGD (DPSGD) optimizer step.
Based on Abadi et al., 2016, this optimizer can be used to reduce the sensitivity of the model to individual training samples or groups thereof, and can thus be used to train models on sensitive data. The optimizer has a number of parameters that are potentially data dependent, and must be provided by the user. IMPORTANT: this optimizer, unlike the others, requires access to the per-example gradients; thus, the gradients should have a leading "batch" dimension. This can be accomplished by using the VectorizedMap node on the gradient pipeline. Like all step nodes, this node only processes gradients, and the resulting updates must be applied manually to the weights (this can be accomplished using the Add node). However, you can also pass it to the StepSolver node which implements the full optimization loop. The learning rate can instead be given as a schedule, by wiring one of the Schedule nodes into the learning_rate_schedule port. More Info... Version 0.2.0
Ports/Properties¶
gradients¶
Gradients to be transformed.
weights¶
Optional current weights.
state¶
Explicit state of the node.
learning_rate_schedule¶
Optional learning rate schedule.
learning_rate¶
Learning rate. A typical choice may be 0.001 here, but this is problem dependent. If a learning rate schedule is provided, this value should be left unspecified.
l2_norm_clip¶
L2 norm clipping value. Maximum l2-norm of the per-example parameter updates. Must be provided.
noise_multiplier¶
Noise multiplier. Ratio of standard deviation to the clipping norm. Must be provided.
randseed¶
Integer random seed. Must be provided.
momentum¶
Optional exponential decay rate for momentum.
nesterov¶
Whether to use Nesterov acceleration.
set_breakpoint¶
Set a breakpoint on this node. If this is enabled, your debugger (if one is attached) will trigger a breakpoint.
metadata¶
User-definable meta-data associated with the node. Usually reserved for technical purposes.