com.microsoft.BiasSoftmax

com.microsoft · ONNX Runtime contrib operator · contrib since_version 1

Description

Computes softmax(data + bias) over the flattened suffix beginning at axis. The required is_inner_broadcast attribute selects how bias rows are reused: consecutive groups for inner broadcast or cyclic groups for outer broadcast. This specializes the softmax(scores + additive_mask) pattern used by transformer attention. Float16 and float32 are supported; the schema's double type is not.

See the ONNX Runtime BiasSoftmax contrib-operator spec for the reference semantics.

Inputs

Name Bind key Logical dtype Rank Shape Description Presence
data data T The input data tensor. required
bias bias T The bias (or additive mask) tensor. Its element count must be an integral number of flattened softmax rows and that row count must divide the data row count. required

Outputs

Name Bind key Logical dtype Rank Shape Description Presence
output output T same as data same as data The output tensor; same shape as data. required

Attributes

Attributes and default values (overridable per request):

Attribute Default Description
axis 1 The axis from which softmax is applied; dimensions from axis onward are included in the softmax reduction.
is_inner_broadcast When 1, bias is broadcast across dimensions from broadcast_axis to axis-1; when 0, bias is broadcast across dimensions 0 to broadcast_axis-1.

Type constraints

Variable Allowed dtypes
T float32, float16

Files

Use with @huggingface/kernels

The loader derives every required output's shape and logical dtype from the manifest contract and this call. It then allocates the result tensors automatically.

The version: 1 option selects the published kernel contract; it is independent of any operator opset, contrib since_version, or model version.

Replace each *Data placeholder with a typed array containing the corresponding input data.

import { getKernel } from "@huggingface/kernels";

const kernel = await getKernel("webgpu-kernels/com.microsoft.BiasSoftmax", { version: 1 });
const { output } = await kernel({
  data: { data: dataData, shape: [1, 2, 2] },
  bias: { data: biasData, shape: [1, 2, 2] },
}, {
  attrs: { is_inner_broadcast: 1 },
});
Downloads last month
-
kernel
webgpu
wgsl
apache-2.0
WebGPU

Requires WebGPU support. See the compatibility table.