Skip to content

← deep_learning package

TrustRatioScalingStep

A chainable step that scales gradients by the trust ratio (ratio of parameter norm to update norm).

This is the underlying raw scaling rule of the Fromage, LARS, and LAMB optimizers, and not an end-to-end optimizer by itself. See also You et al. (2020) for an analysis. More Info... Version 0.2.0

Ports/Properties

gradients

Gradients to be transformed.

verbose name
Gradients
default value
None
port type
DataPort
value type
object (can be None)
data direction
INOUT

weights

Optional current weights.

verbose name
Weights
default value
None
port type
DataPort
value type
object (can be None)
data direction
IN

state

Explicit state of the node.

verbose name
State
default value
None
port type
DataPort
value type
object (can be None)
data direction
INOUT

min_norm

Minimum gradient norm. This can be used to avoid dividing by zero when rescaling; small gradients are rescaled to at least this value.

verbose name
Min Norm
default value
0
port type
FloatPort
value type
float (can be None)

trust_coefficient

Trust coefficient. A multiplier applied to the trust ratio, can be used to scale the update size.

verbose name
Trust Coefficient
default value
1
port type
FloatPort
value type
float (can be None)

epsilon

Small value applied to the denominator outside to avoid dividing by zero when rescaling.

verbose name
Epsilon
default value
0
port type
FloatPort
value type
float (can be None)

set_breakpoint

Set a breakpoint on this node. If this is enabled, your debugger (if one is attached) will trigger a breakpoint.

verbose name
Set Breakpoint (Debug Only)
default value
False
port type
BoolPort
value type
bool (can be None)

metadata

User-definable meta-data associated with the node. Usually reserved for technical purposes.

verbose name
Metadata
default value
{}
port type
DictPort
value type
dict (can be None)