RAdamStep¶
The Rectified Adam optimizer step.
Based on Liu et al., 2020, Rectified Adam addresses a shortcoming in the popular Adam optimizer, where during initial stages of training, the gradients exhibit a large variance due to the limited number of training samples used to estimate the optimizer's statistics, which typically is addressed using warm-up schedules. Like all step nodes, this node only processes gradients, and the resulting updates must be applied manually to the weights (this can be accomplished using the Add node). However, you can also pass it to the StepSolver node which implements the full optimization loop. More Info... Version 0.2.0
Ports/Properties¶
gradients¶
Gradients to be transformed.
weights¶
Optional current weights.
state¶
Explicit state of the node.
learning_rate_schedule¶
Optional learning rate schedule.
learning_rate¶
Learning rate. A typical choice may be 0.001 here, but this is problem dependent. If a learning rate schedule is provided, this value should be left unspecified.
beta1¶
Exponential decay rate for the first moment estimates.
beta2¶
Exponential decay rate for the second moment estimates.
epsilon¶
Small value applied to the denominator outside the square root to avoid dividing by zero when rescaling.
epsilon_inroot¶
Small value applied to the denominator inside the square root to avoid dividing by zero when rescaling. A case where this is needed is when differentiating the optimizer itself, eg for bilevel optimization.
threshold¶
Threshold for variance tractability.
set_breakpoint¶
Set a breakpoint on this node. If this is enabled, your debugger (if one is attached) will trigger a breakpoint.
metadata¶
User-definable meta-data associated with the node. Usually reserved for technical purposes.