Back to Top

gpu.js 2.24.1

GPU Accelerated JavaScript

alias(name, source)

Parameters

Name Type Description
name
source

Returns

Function

createKernel()

A kernel may also be given as source text. This is the only form available where the engine does not retain function source — React Native's Hermes, for instance, returns "function name(a0, a1) { [bytecode] }" from Function.prototype.toString().

Returns

Void

createPipeline()

Compile a whole multi-kernel computation into one callable plan. The orchestration function runs once, at build time (first call), with opaque handles for arguments; the kernel calls it makes are recorded and replayed on later calls with intermediates kept resident. Calling the pipeline always returns a Promise.

Returns

Void

fallbackReason()

why this kernel's work was degraded to the cpu backend, when it was (#868)

Returns

Void

moduleCacheLimit()

LRU bound on cached per-size-signature wasm instantiations (#870)

Returns

Void

poolSize()

worker-pool size cap for threaded runs; null lets the pool decide

Returns

Void

kernelOrder()

webasm sits last, one step above the cpu fallback: any working GL backend outranks it, so auto modes only reach it where no GL context exists

Returns

Void

kernelTypes()

Returns

Void

new GPU()

The GPU.js library class which manages the GPU context for the creating kernels

Returns

GPU

isWebGPUAvailable()

Returns

Promise.<boolean>

constructor([settings])

Creates an instance of GPU.

Parameters

Name Type Description
settings IGPUSettings
  • Settings to set mode, and other properties
Optional

Returns

Void

this._webGPUDecision()

mode 'async' only: the adapter probe's settled answer (true/false), or null while it is in flight. Started at construction so that by kernel creation -- usually at least a task later in real applications -- the backend for a graphical kernel can be decided BEFORE its canvas is exposed, since a canvas is permanently committed to its first context type and can never be swapped between backends afterwards.

Returns

Void

chooseKernel()

Choose kernel type and save on .Kernel property of GPU

Returns

Void

Kernel()

Returns

Void

createKernel(source[, settings])

Parameters

Name Type Description
source Function String object
  • The calling to perform the conversion
settings IGPUKernelSettings
  • The parameter configuration object
Optional

Returns

IKernelRunShortcut

callable function to run

onRequestSwitchKernel(reasons, args, _kernel)

Parameters

Name Type Description
reasons Array.<IReason>
args IArguments
_kernel Kernel

Returns

createPipeline(fn[, settings])

Parameters

Name Type Description
fn Function
  • orchestration function; may only call kernels created by this GPU instance
settings IPipelineSettings
  • constants only in v1
Optional

Returns

IPipelineRunShortcut

callable pipeline

createKernelMap(subKernels, rootKernel)

Create a super kernel which executes sub kernels and saves their output to be used with the next sub kernel. This can be useful if we want to save the output on one kernel, and then use it as an input to another kernel. Machine Learning

Parameters

Name Type Description
subKernels Object Array
  • Sub kernels for this kernel
rootKernel Function
  • Root kernel

Examples

const megaKernel = gpu.createKernelMap({
  addResult: function add(a, b) {
    return a[this.thread.x] + b[this.thread.x];
  },
  multiplyResult: function multiply(a, b) {
    return a[this.thread.x] * b[this.thread.x];
  },
 }, function(a, b, c) {
      return multiply(add(a, b), c);
});

megaKernel(a, b, c);

Note: You can also define subKernels as an array of functions.
> [add, multiply]

Returns

Function

callable kernel function

combineKernels(subKernels, rootKernel)

Combine different kernels into one super Kernel, useful to perform multiple operations inside one kernel without the penalty of data transfer between cpu and gpu.

The number of kernel functions sent to this method can be variable. You can send in one, two, etc.

Parameters

Name Type Description
subKernels Function
  • Kernel function(s) to combine.
rootKernel Function
  • Root kernel to combine kernels into

Examples

  combineKernels(add, multiply, function(a,b,c){
    return add(multiply(a,b), c)
 })

Returns

Function

Callable kernel function

addFunction(source[, settings])

Parameters

Name Type Description
source Function String
  • Javascript function to convert
settings IFunctionSettings Optional

Returns

GPU

returns itself

addNativeFunction(name, source[, settings])

Parameters

Name Type Description
name String
  • native function name, used for reverse lookup
source String
  • the native function implementation, as it would be defined in it's entirety
settings object Optional

Returns

GPU

returns itself

injectNative(source)

Inject a string just before translated kernel functions

Parameters

Name Type Description
source String

Returns

GPU

destroy()

Returns

Promise

kernelRunShortcut(kernel)

Makes kernels easier for mortals (including me)

Parameters

Name Type Description
kernel

Returns

function()

shortcut.exec()

Run kernel in async mode

Returns

Promise.<KernelOutput>

MSG_HANDLE_READ()

Pipeline compilation (docs/design/pipeline-compilation.md): the orchestration function runs ONCE, at build time, against opaque handles; every kernel call made while the trace is open is recorded into a static plan, and later pipeline calls execute the plan without re-entering user code. JS loops in the orchestration therefore unroll at trace time, and closure-captured plain values freeze into the plan the same way constants do.

Returns

Void

new PipelineHandle()

The class exists for instanceof and for its name in errors; all state lives in the trace's WeakMap so the frozen instance has no own properties for the Proxy get trap to conflict with.

Returns

Void

activeTrace()

Consulted by kernelRunShortcut on every call; non-null only while an orchestration function is being traced, which is always synchronous, so a module-level slot cannot see two traces at once.

Returns

Void

new PipelineTrace()

Trace-time state: records kernel calls as plan steps and mints the opaque handles that stand in for values the orchestration never gets to see.

Returns

Void

this.kernels()

distinct kernel run-shortcuts, in first-use order; steps refer to them by index so the ping-pong loop shape compiles to ONE kernel entry

Returns

Void

createHandle(meta)

Parameters

Name Type Description
meta Object
  • {source: 'pipelineArg', index} | {source: 'step', step}

Returns

Proxy.<PipelineHandle>

recordKernelCall(shortcut, args)

Entry point from kernelRunShortcut while a trace is open: validate the kernel, bind the arguments, and answer with a fresh step-output handle instead of running anything.

Parameters

Name Type Description
shortcut IKernelRunShortcut
args IArguments

Returns

Proxy.<PipelineHandle>

bindValue()

Returns

Object

argBinding per the plan IR; non-handles snapshot here, which is the moment closure-captured mutables freeze

snapshotValue()

Call-time sampling: mutable JS values copy before the call promise can yield, so const p = pipeline(buf); buf[0] = 9; computes on the value buf held at the call. Handles never reach this function -- bindValue checks the WeakMap first -- so property access here cannot trip a handle trap.

Returns

Void

assignBuffers(steps, resultBindings)

Static liveness over the unrolled DAG, then greedy slot reuse: a step may write a buffer only when the previous occupant's last reader ran strictly earlier -- a reader AT the writing step still needs the old contents while the new ones are produced, which is exactly what forces u = sweep(u, q) in a loop onto two alternating buffers. Slots are only shared between steps of identical output shape so the fused executor can lay them out as fixed regions.

Parameters

Name Type Description
steps Array
  • mutated: outputBuffer assigned per step
resultBindings Array

Returns

Array

buffers

bindResults(trace, returned)

Parameters

Name Type Description
trace PipelineTrace
returned
  • the orchestration function's return value

Returns

Object

results descriptor {kind, entries: [{key?, binding}]}

constructor(gpu, fn[, settings])

Parameters

Name Type Description
gpu GPU
fn Function
  • orchestration function, run once per (re)trace
settings IPipelineSettings Optional

Returns

Void

this.executorKind()

executor identity probe for tests and later phases: 'generic' executes step-by-step through the normal kernel machinery on every backend; 'fused-sync' is the webasm executor running every step over one shared wasm memory; 'fused-threaded' is that executor with pool workers walking the whole plan on an Atomics barrier; 'fused-encoder' is the webgpu executor recording every step into one command encoder over persistent storage buffers

Returns

Void

this.fallbackReason()

why the fused executor declined this plan; null while fused (or before the first call decides)

Returns

Void

this._executor()

undefined: not yet attempted for this plan; false: attempted and declined (generic runs); otherwise the compiled fused executor

Returns

Void

this._fusionDisabled()

test/benchmark hook: forces the generic executor when true

Returns

Void

this.destroyed()

test/benchmark hook: keeps a fused executor off the worker pool

Returns

Void

this._tail()

concurrent calls to one pipeline serialize on this tail, the same contract as threaded webasm kernels

Returns

Void

call(args)

Parameters

Name Type Description
args IArguments Array

Returns

Promise.<*>

_guardAsync(result)

The threaded executor rejects asynchronously (worker death, stalled barrier, destroy mid-run); any such failure leaves its barrier state unusable, so the executor is dropped and the next call compiles a fresh one. Fallback decisions stay synchronous — the signature check throws before dispatch — so a FusionFallback can never surface here.

Parameters

Name Type Description
result
  • executor.execute's return: a value (sync) or a Promise (threaded)

Returns

Void

setConstants(constants)

Parameters

Name Type Description
constants Object

Returns

Pipeline

destroy()

Returns

Promise

_buildPlan()

Runs the orchestration function once with handles for arguments; the recorded steps become the plan. Math.random is barred for the duration because a trace-time draw would freeze into every later call.

Returns

Object

plan IR

_genericClone(plan, step)

The generic executor's writer for one (kernel, seat signature, output slot) triple. One MUTABLE, STATICALLY-TYPED clone per triple reproduces the hand-rolled two-kernel ping-pong mechanically: each clone owns one output texture/array for the life of the plan (steady state allocates nothing per step) and sees one argument-type signature (no per-call dynamicArguments re-typing -- the forced re-typing was most of a 7x loss even after the texture churn was gone). Static liveness (assignBuffers) is what makes mutability safe: no step ever reads a slot while that slot's writer renders. Argument drift across CALLS is the clones' own switch machinery's business, as for any kernel.

Parameters

Name Type Description
plan Object
step Object

Returns

IKernelRunShortcut

_prepareExecutor(args)

Attempts the backend's fused executor for the current plan against this call's sampled arguments: the single-encoder lowering on webgpu (async — kernel builds await the device), the wasm-memory lowering everywhere else. Anything the fused compile cannot take degrades to the generic executor with the reason recorded, its usual degradation contract.

Parameters

Name Type Description
args Array
  • sampled pipeline arguments; sizes/types bake into the fused layout

Returns

Promise

_cloneKernel(shortcut)

The plan runs on private instances configured for pipeline use -- pipeline: true, immutable: true -- so intermediates stay resident (textures on GL, fresh arrays on cpu) and the user's kernel settings are never observably touched. Kernels stay shared between pipelines and direct use through their own shortcuts.

Parameters

Name Type Description
shortcut IKernelRunShortcut
  • the user's kernel

Returns

IKernelRunShortcut

private clone

_uploadArg()

Array pipeline arguments upload ONCE per call on backends where an upload costs (GL textures, webgpu buffers): a lazy per-arg identity kernel parks the value device-side and every consuming step binds the handle -- feeding the raw array to a 200-step plan re-uploaded it 200 times, which was most of the remaining gap to hand-rolled ping-pong. cpu/webasm consume arrays natively, so there the raw value is optimal.

Returns

Void

_genericEagerUploadsPay()

Eager uploads are only sound where the upload call is SYNCHRONOUS (the GL family): the texture materializes before user code can run again. webgpu uploads return promises, so its generic path keeps copies.

Returns

Void

utils()

Returns

Void

systemEndianness()

Returns

String

'LE' or 'BE' depending on system architecture Credit: https://gist.github.com/TooTallNate/4750953

isFunction(funcObj)

Parameters

Name Type Description
funcObj Function
  • Object to validate if its a function

Returns

Boolean

TRUE if the object is a JS function

isFunctionString(fn)

Parameters

Name Type Description
fn String
  • String of JS function to validate

Returns

Boolean

TRUE if the string passes basic validation

getFunctionNameFromString(funcStr)

Parameters

Name Type Description
funcStr String
  • String of JS function to validate

Returns

String

Function name string (if found)

getArgumentNamesFromString(fn)

Parameters

Name Type Description
fn String
  • String of JS function to validate

Returns

Array.<String>

Array representing all the parameter names

clone(obj)

Parameters

Name Type Description
obj Object
  • Object to clone

Returns

Object Array

Cloned object

isArray(array)

Parameters

Name Type Description
array Object
  • The argument object to check if is array

Returns

Boolean

true if is array or Array-like object

typeFitsValue(type, value)

Parameters

Name Type Description
type String
value

Returns

Boolean

closestSquareDimensions(length)

Parameters

Name Type Description
length Number

Returns

TextureDimensions

getMemoryOptimizedFloatTextureSize(dimensions, bitRatio)

A texture takes up four

Parameters

Name Type Description
dimensions OutputDimensions
bitRatio Number

Returns

TextureDimensions

getMemoryOptimizedPackedTextureSize(dimensions, bitRatio)

Parameters

Name Type Description
dimensions
bitRatio

Returns

TextureDimensions

getDimensions(x[, pad])

Parameters

Name Type Description
x Array String Texture Input
  • The array
pad Boolean
  • To include padding in the dimension calculation
Optional

Returns

OutputDimensions

flatten2dArrayTo(array, target)

Puts a nested 2d array into a one-dimensional target array

Parameters

Name Type Description
array Array
target Float32Array Float64Array

Returns

Void

flatten3dArrayTo(array, target)

Puts a nested 3d array into a one-dimensional target array

Parameters

Name Type Description
array Array
target Float32Array Float64Array

Returns

Void

flatten4dArrayTo(array, target)

Puts a nested 4d array into a one-dimensional target array

Parameters

Name Type Description
array Array
target Float32Array Float64Array

Returns

Void

flattenTo(array, target)

Puts a nested 1d, 2d, or 3d array into a one-dimensional target array

Parameters

Name Type Description
array Float32Array Uint16Array Uint8Array
target Float32Array

Returns

Void

splitArray(array, part)

Parameters

Name Type Description
array Array.<Number>
  • The array to split into chunks
part Number
  • elements in one chunk

Returns

Array.<Number>

An array of smaller chunks

glslFloatLiteral()

A number as a GLSL float literal. Integer-valued numbers at 1e21 and beyond stringify in exponential form, which is already a valid GLSL float literal — appending .0 to it is not (#864).

Returns

Void

linesToString(lines)

Parameters

Name Type Description
lines Array
  • An Array of strings

Returns

String

Single combined String, separated by \n

flattenFunctionToString(source, settings)

Parameters

Name Type Description
source String
settings Object

Returns

String

splitHTMLImageToRGB(gpu, image)

Parameters

Name Type Description
gpu GPU
image

Returns

Array

splitRGBAToCanvases(gpu, rgba, width, height)

A visual debug utility

Parameters

Name Type Description
gpu GPU
rgba
width
height

Returns

Array.<Object>

new Texture(settings)

Parameters

Name Type Description
settings IGPUTextureSettings

Returns

Void

this.kernel()

Returns

Void

toArray()

Returns

TextureArrayOutput

clone()

Returns

Texture

delete()

Returns

Void

new FunctionBuilder()

Returns

Void

FunctionBuilder.fromKernel(kernel, FunctionNode[, extraNodeOptions])

Parameters

Name Type Description
kernel Kernel
FunctionNode FunctionNode
extraNodeOptions object Optional

Returns

FunctionBuilder

FunctionBuilder.constructor([settings])

Parameters

Name Type Description
settings IFunctionBuilderSettings Optional

Returns

Void

FunctionBuilder.addFunctionNode(functionNode)

Parameters

Name Type Description
functionNode FunctionNode
  • functionNode to add

Returns

Void

FunctionBuilder.traceFunctionCalls(functionName[, retList])

Parameters

Name Type Description
functionName String
  • Function name to trace from, default to 'kernel'
retList Array.<String>
  • Returning list of function names that is traced. Including itself.
Optional

Returns

Array.<String>

Returning list of function names that is traced. Including itself.

dependantNativeFunctionName()

https://github.com/gpujs/gpu.js/issues/207 if dependent function is already in the list, because a function depends on it, and because it has already been traced, we know that we must move the dependent function to the end of the the retList.

Returns

Void

dependantFunctionName()

https://github.com/gpujs/gpu.js/issues/207 if dependent function is already in the list, because a function depends on it, and because it has already been traced, we know that we must move the dependent function to the end of the the retList.

Returns

Void

getPrototypeString(functionName)

Parameters

Name Type Description
functionName String
  • Function name to trace from. If null, it returns the WHOLE builder stack

Returns

String

The full string, of all the various functions. Trace optimized if functionName given

getPrototypes([functionName])

Parameters

Name Type Description
functionName String
  • Function name to trace from. If null, it returns the WHOLE builder stack
Optional

Returns

Array

The full string, of all the various functions. Trace optimized if functionName given

getStringFromFunctionNames(functionList)

Parameters

Name Type Description
functionList Array.<String>
  • List of function to build string

Returns

String

The string, of all the various functions. Trace optimized if functionName given

getPrototypesFromFunctionNames(functionList)

Parameters

Name Type Description
functionList Array.<String>
  • List of function names to build the string.

Returns

Array

Prototypes of all functions converted

getString(functionName)

Parameters

Name Type Description
functionName String
  • Function name to trace from. If null, it returns the WHOLE builder stack

Returns

String

settings - The string, of all the various functions. Trace optimized if functionName given

_getFunction(functionName) private method

Parameters

Name Type Description
functionName String

Returns

FunctionNode

lookupFunctionArgumentBitRatio(functionName, argumentName)

Parameters

Name Type Description
functionName string
argumentName string

Returns

number

assignArgumentBitRatio(functionName, argumentName, calleeFunctionName, argumentIndex)

Parameters

Name Type Description
functionName string
argumentName string
calleeFunctionName string
argumentIndex number

Returns

number

new FunctionNode()

Returns

Void

FunctionNode.constructor(source[, settings])

Parameters

Name Type Description
source string object
settings IFunctionSettings Optional

Returns

Void

FunctionNode.isIdentifierConstant(name)

Parameters

Name Type Description
name String

Returns

boolean

FunctionNode.astMemberExpressionUnroll(ast)

Parameters

Name Type Description
ast Object
  • the AST object to parse

Returns

String

the function namespace call, unrolled

FunctionNode.requiresSequenceFreeForInit()

Whether this backend needs for (a, b; ...) inits hoisted to statements before the loop -- WGSL cannot express the comma; GLSL and JS take it natively.

Returns

Boolean

FunctionNode.getAssignedArguments()

Returns

Set.<String>

original (unsanitized) argument names

FunctionNode.getVariableType(ast)

Parameters

Name Type Description
ast Object
  • Identifier

Returns

String

Type of the parameter

FunctionNode.getLookupType(type)

Generally used to lookup the value type returned from a member expressions

Parameters

Name Type Description
type String

Returns

String

FunctionNode.getType(ast)

Recursively looks up type for ast expression until it's found

Parameters

Name Type Description
ast

Returns

String

FunctionNode.getDependencies(ast, dependencies, isNotSafe)

Parameters

Name Type Description
ast
dependencies
isNotSafe

Returns

Array

FunctionNode.astGeneric(ast, retArr)

Parameters

Name Type Description
ast Object
  • the AST object to parse
retArr Array
  • return array string

Returns

Array

the parsed string array

FunctionNode.astErrorOutput(error, ast)

Parameters

Name Type Description
error string
  • the error message output
ast Object
  • the AST object where the error is

Returns

Void

FunctionNode.astFunction(ast, retArr)

Parameters

Name Type Description
ast Object
retArr Array.<String>

Returns

Array.<String>

FunctionNode.astFunctionDeclaration(ast, retArr)

Parameters

Name Type Description
ast Object
  • the AST object to parse
retArr Array
  • return array string

Returns

Array

the append retArr

FunctionNode.astExpressionStatement(esNode, retArr)

Parameters

Name Type Description
esNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the append retArr

FunctionNode.astEmptyStatement(eNode, retArr)

Parameters

Name Type Description
eNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the append retArr

FunctionNode.astBreakStatement(brNode, retArr)

Parameters

Name Type Description
brNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the append retArr

FunctionNode.astContinueStatement(crNode, retArr)

Parameters

Name Type Description
crNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the append retArr

FunctionNode.astVariableDeclarator(iVarDecNode, retArr)

Parameters

Name Type Description
iVarDecNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the append retArr

FunctionNode.astUnaryExpression(uNode, retArr)

Parameters

Name Type Description
uNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the append retArr

FunctionNode.astUpdateExpression(uNode, retArr)

Parameters

Name Type Description
uNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the append retArr

FunctionNode.astLogicalExpression(logNode, retArr)

Parameters

Name Type Description
logNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the append retArr

FunctionNode.getMemberExpressionDetails(ast)

Parameters

Name Type Description
ast

Returns

IFunctionNodeMemberExpressionDetails

minifiedSyntheticId()

De-minification: minifiers (esbuild, terser) fold statements into expressions -- if (c) { x = 1; } becomes c && (x = 1), statement sequences become comma expressions, if/else becomes a ternary of assignments. In statement position the folded expression's VALUE is discarded, so unfolding back into statements is always semantics-preserving, no side-effect analysis required. Runs on the parsed AST before FunctionTracer records anything, so every backend sees plain statements; on webgl it also runs before the FXC hoisting normalization, which only understands statement shapes.

Synthetic nodes are stamped with unique start/end: astKey and the literal-type cache are keyed by position. The 0x20000000 base is disjoint from real acorn offsets and from the 0x40000000 base the webgl hoisting machinery stamps its own synthetic nodes with.

Returns

Void

normalizeMinifiedNested()

A single-statement position (an unbraced loop body or if branch) that unfolds into several statements needs a block around them.

Returns

Void

normalizeMinifiedForHeader()

Loop simplification for minified for-headers. for (i = 0, j = 0; test; i++, j++) cannot be expressed on every backend (WGSL takes one statement per clause), so a comma INIT hoists to statements before the loop and a comma UPDATE moves to the end of the body -- with a copy ahead of every continue that belongs to this loop, preserving per-iteration timing. Returns the statements to place before the loop. A labeled continue makes the update rewrite unsafe, so such a loop is left exactly as written.

Returns

Void

cloneWithSyntheticPositions()

A deep copy with fresh synthetic positions on every node: the same update lands both at the body's end and ahead of each continue, and position-keyed caches must see distinct nodes.

Returns

Void

prependBeforeContinues()

Puts a copy of prefix ahead of every continue belonging to this loop. Nested loops keep their own continues. Returns the rewritten block, or null when a labeled continue makes the rewrite unsafe.

Returns

Void

this.declarations()

Returns

Void

getDeclaration(name)

Parameters

Name Type Description
name string

Returns

IDeclaration

scan(ast)

Recursively scans AST for declarations and functions, and add them to their respective context

Parameters

Name Type Description
ast

Returns

Void

isContextMatch()

Returns

Boolean

getFeatures()

Returns

Void

constructor(source, settings)

Parameters

Name Type Description
source string IKernelJSON
settings

Returns

Void

this.onRequestSwitchKernel()

Supplied by GPU.createKernel; swaps in a kernel compiled for the arguments this one was handed. Declared here rather than on the GL kernel so mergeSettings carries it onto every backend -- the cpu kernel needs it for the same argument-type changes.

Returns

Void

this.argumentNames()

Name of the arguments found from parsing source argument

Returns

Void

this.source()

The function source

Returns

Void

this.output()

The size of the kernel's output

Returns

Void

this.debug()

Debug mode

Returns

Void

this.graphical()

Graphical mode

Returns

Void

this.loopMaxIterations()

Maximum loops when using argument values to prevent infinity

Returns

Void

this.constants()

Constants used in kernel via this.constants

Returns

Void

this.constantTypes()

Returns

Void

this.constantBitRatios()

Returns

Void

this.dynamicArguments()

Returns

Void

this.dynamicOutput()

Returns

Void

this.canvas()

Returns

Void

this.context()

Returns

Void

this.checkContext()

Returns

Void

this.gpu()

Returns

Void

this.functions()

Returns

Void

this.nativeFunctions()

Returns

Void

this.injectedNative()

Returns

Void

this.subKernels()

Returns

Void

this.validate()

Returns

Void

this.immutable()

Enforces kernel to write to a new array or texture on run

Returns

Void

this.pipeline()

Enforces kernel to write to a texture on run

Returns

Void

this.asyncMode()

Makes the kernel return a Promise of its result on every backend. Backends with a genuinely non-blocking readback (webgl2 fences, webgpu natively) use it; the rest resolve their synchronous result, so the calling contract is uniform either way.

Returns

Void

this.precision()

Make GPU use single precision or unsigned. Acceptable values: 'single' or 'unsigned'

Returns

Void

this.tactic()

Returns

Void

this.randomSeed()

Seed for Math.random() so kernel runs are reproducible; null seeds from Math.random()

Returns

Void

this.switchingKernels()

Reasons this kernel cannot serve the call it was handed, collected for the caller's switch; null when it can.

Returns

Void

mergeSettings(settings)

Parameters

Name Type Description
settings IDirectKernelSettings IJSONSettings

Returns

Void

build()

Returns

Void

run()

Returns

Float32Array Array.<Float32Array> Array.<Array.<Float32Array>>

Result The final output of the program, as float, and as Textures for reuse.

initCanvas()

Returns

Object

initContext()

Returns

Object

initPlugins(settings)

Parameters

Name Type Description
settings IDirectKernelSettings

Returns

{string[]};

addFunction(source[, settings])

Parameters

Name Type Description
source KernelFunction string IGPUFunction
settings IFunctionSettings Optional

Returns

Kernel

addNativeFunction(name, source[, settings])

Parameters

Name Type Description
name string
source string
settings IGPUFunctionSettings Optional

Returns

Void

setupArguments(args)

Parameters

Name Type Description
args IArguments
  • The actual parameters sent to the Kernel

Returns

Void

setupConstants()

Setup constants

Returns

Void

setOptimizeFloatMemory(flag)

Parameters

Name Type Description
flag

Returns

this

toKernelOutput(output)

Parameters

Name Type Description
output Array Object

Returns

Array.<number>

setOutput(output)

Parameters

Name Type Description
output Array Object
  • The output array to set the kernel output size to

Returns

this

setDebug(flag)

Parameters

Name Type Description
flag Boolean
  • true to enable debug

Returns

this

setGraphical(flag)

Parameters

Name Type Description
flag Boolean
  • true to enable graphical output

Returns

this

setLoopMaxIterations(max)

Parameters

Name Type Description
max number
  • iterations count

Returns

this

setConstants()

Returns

this

setConstantTypes(constantTypes)

Parameters

Name Type Description
constantTypes IKernelValueTypes

Returns

this

setFunctions(functions)

Parameters

Name Type Description
functions Array.<IFunction> Array.<KernelFunction>

Returns

this

setNativeFunctions(nativeFunctions)

Parameters

Name Type Description
nativeFunctions Array.<IGPUNativeFunction>

Returns

this

setInjectedNative(injectedNative)

Parameters

Name Type Description
injectedNative String

Returns

this

setPipeline(flag)

Set writing to texture on/off

Parameters

Name Type Description
flag

Returns

this

setAsyncMode(flag)

Set Promise-returning mode on/off

Parameters

Name Type Description
flag Boolean

Returns

this

setPrecision(flag)

Set precision to 'unsigned' or 'single'

Parameters

Name Type Description
flag String

'unsigned' or 'single'

Returns

this

setDimensions(flag)

Parameters

Name Type Description
flag

Returns

Kernel

setOutputToTexture(flag)

Parameters

Name Type Description
flag

Returns

this

setImmutable(flag)

Set to immutable

Parameters

Name Type Description
flag

Returns

this

setCanvas(canvas)

Parameters

Name Type Description
canvas Object

Returns

this

setStrictIntegers(flag)

Parameters

Name Type Description
flag Boolean

Returns

this

setDynamicOutput(flag)

Parameters

Name Type Description
flag

Returns

this

setRandomSeed(seed)

Set a seed for Math.random(), so kernel runs are reproducible

Parameters

Name Type Description
seed Number

Returns

this

setHardcodeConstants(flag)

Parameters

Name Type Description
flag

Returns

this

setDynamicArguments(flag)

Parameters

Name Type Description
flag

Returns

this

setUseLegacyEncoder(flag)

Parameters

Name Type Description
flag Boolean

Returns

this

setWarnVarUsage(flag)

Parameters

Name Type Description
flag Boolean

Returns

this

getCanvas()

Returns

Object

getWebGl()

Returns

Object

setContext(context)

Parameters

Name Type Description
context WebGLRenderingContext
  • webGl instance to bind

Returns

Void

setArgumentTypes(argumentTypes)

Parameters

Name Type Description
argumentTypes IKernelValueTypes Array.<GPUVariableType>

Returns

this

setTactic(tactic)

Parameters

Name Type Description
tactic Tactic

Returns

this

requestFallback(args[, reason])

Parameters

Name Type Description
args IArguments
reason String
  • why this kernel cannot run here; carried onto the replacement kernel as fallbackReason and named in the console warning, so the degradation is discoverable (#868)
Optional

Returns

Void

validateSettings()

Returns

Void

addSubKernel(subKernel)

Parameters

Name Type Description
subKernel ISubKernel
  • function (as a String) of the subKernel to add

Returns

Void

destroy([removeCanvasReferences])

Parameters

Name Type Description
removeCanvasReferences Boolean

remove any associated canvas references

Optional

Returns

Void

getBitRatio(value)

bit storage ratio of source to target 'buffer', i.e. if 8bit array -> 32bit tex = 4

Parameters

Name Type Description
value

Returns

number

getPixels([flip])

Parameters

Name Type Description
flip Boolean Optional

Returns

Uint8ClampedArray

prependString(value)

Parameters

Name Type Description
value String

Returns

Void

hasPrependString(value)

Parameters

Name Type Description
value String

Returns

Boolean

toJSON()

Returns

IKernelJSON

buildSignature(args)

Parameters

Name Type Description
args IArguments

Returns

Void

getArgumentTypes(kernel, args)

Parameters

Name Type Description
kernel Kernel
args IArguments

Returns

GPUVariableType[]

getSignature(kernel, argumentTypes)

Parameters

Name Type Description
kernel Kernel
argumentTypes Array.<GPUVariableType>

Returns

Void

functionToIGPUFunction(source[, settings])

Parameters

Name Type Description
source String Function
settings IFunctionSettings Optional

Returns

IGPUFunction

onActivate(previousKernel)

Parameters

Name Type Description
previousKernel Kernel

Returns

Void

switchKernels(reason)

Parameters

Name Type Description
reason IReason

Returns

Void

checkArgumentTypes(args)

Parameters

Name Type Description
args IArguments Array

Returns

Void

new KernelValue()

Returns

Void

KernelValue.constructor(value, settings)

Parameters

Name Type Description
value KernelVariable
settings IKernelValueSettings

Returns

Void

module.exports()

Returns

Void

mulberry32(seed)

mulberry32, a fast counter-based PRNG with good distribution

Parameters

Name Type Description
seed Number

Returns

Function

onBeforeRun(kernel)

Parameters

Name Type Description
kernel Kernel

Returns

Void

plugin()

Returns

Void

glWiretap(gl[, options])

Parameters

Name Type Description
gl WebGLRenderingContext
options IGLWiretapOptions Optional

Returns

GLWiretapProxy

glExtensionWiretap(extension, options)

Parameters

Name Type Description
extension
options IGLExtensionWiretapOptions

Returns

new CPUFunctionNode()

Returns

Void

CPUFunctionNode.markupUserName()

Returns

Void

CPUFunctionNode.astReturnStatement(ast, retArr)

Parameters

Name Type Description
ast Object
  • the AST object to parse
retArr Array
  • return array string

Returns

Array

the append retArr

CPUFunctionNode.astLiteral(ast, retArr)

Parameters

Name Type Description
ast Object
  • the AST object to parse
retArr Array
  • return array string

Returns

Array

the append retArr

CPUFunctionNode.astBinaryExpression(ast, retArr)

Parameters

Name Type Description
ast Object
  • the AST object to parse
retArr Array
  • return array string

Returns

Array

the append retArr

CPUFunctionNode.astIdentifierExpression(idtNode, retArr)

Parameters

Name Type Description
idtNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the append retArr

CPUFunctionNode.astForStatement(forNode, retArr)

Parameters

Name Type Description
forNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the parsed webgl string

CPUFunctionNode.astWhileStatement(whileNode, retArr)

Parameters

Name Type Description
whileNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the parsed javascript string

CPUFunctionNode.astDoWhileStatement(doWhileNode, retArr)

Parameters

Name Type Description
doWhileNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the parsed webgl string

CPUFunctionNode.astAssignmentExpression(assNode, retArr)

Parameters

Name Type Description
assNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the append retArr

CPUFunctionNode.astBlockStatement(bNode, retArr)

Parameters

Name Type Description
bNode Object
  • the AST object to parse
retArr Array
  • return array string

Returns

Array

the append retArr

CPUFunctionNode.astVariableDeclaration(varDecNode, retArr)

Parameters

Name Type Description
varDecNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the append retArr

CPUFunctionNode.astIfStatement(ifNode, retArr)

Parameters

Name Type Description
ifNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the append retArr

CPUFunctionNode.astThisExpression(tNode, retArr)

Parameters

Name Type Description
tNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the append retArr

CPUFunctionNode.astMemberExpression(mNode, retArr)

Parameters

Name Type Description
mNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the append retArr

CPUFunctionNode.astCallExpression(ast, retArr)

Parameters

Name Type Description
ast Object
  • the AST object to parse
retArr Array
  • return array string

Returns

Array

the append retArr

CPUFunctionNode.astArrayExpression(arrNode, retArr)

Parameters

Name Type Description
arrNode Object
  • the AST object to parse
retArr Array
  • return array string

Returns

Array

the append retArr

toString()

Returns

Void

new CPUKernel()

Returns

Void

validateSettings(args)

Parameters

Name Type Description
args IArguments

Returns

Void

build()

Returns

Void

getKernelString()

Returns

String

result

toString()

Returns

Void

_getLoopMaxString()

Returns

String

result

getPixels(flip)

Parameters

Name Type Description
flip

Returns

Uint8ClampedArray

glKernelString(Kernel, args, originKernel[, setupContextString, destroyContextString])

Parameters

Name Type Description
Kernel GLKernel
args Array.<KernelVariable>
originKernel Kernel
setupContextString string Optional
destroyContextString string Optional

Returns

string

findKernelValue(argument, kernelValues, values, context, uploadedValues)

Parameters

Name Type Description
argument KernelVariable
kernelValues Array.<KernelValue>
values Array.<KernelVariable>
context
uploadedValues Array.<KernelVariable>

Returns

string

new GLKernel()

Returns

Void

setupFeatureChecks()

Returns

Void

setFixIntegerDivisionAccuracy(fix)

Parameters

Name Type Description
fix Boolean
  • should fix

Returns

Void

setPrecision(flag)

Parameters

Name Type Description
flag String
  • 'single' or 'unsigned'

Returns

Void

setFloatTextures(flag)

Parameters

Name Type Description
flag Boolean
  • true to enable floatTextures

Returns

Void

nativeFunctionArguments(source)

A highly readable very forgiving micro-parser for a glsl function that gets argument types

Parameters

Name Type Description
source String

Returns

[object Object]

this.TextureConstructor()

Returns

Void

pickRenderStrategy(args)

Picks a render strategy for the now finally parsed kernel

Parameters

Name Type Description
args

Returns

KernelOutput

getKernelString()

Returns

String

getMainResultKernelNumberTexture()

Returns

String[]

getMainResultSubKernelNumberTexture()

Returns

String[]

getMainResultKernelArray2Texture()

Returns

String[]

getMainResultSubKernelArray2Texture()

Returns

String[]

getMainResultKernelArray3Texture()

Returns

String[]

getMainResultSubKernelArray3Texture()

Returns

String[]

getMainResultKernelArray4Texture()

Returns

String[]

getMainResultSubKernelArray4Texture()

Returns

String[]

getMainResultGraphical()

Returns

String[]

getMainResultMemoryOptimizedFloats()

Returns

String[]

getMainResultPackedPixels()

Returns

String[]

getFloatTacticDeclaration()

Returns

string

getIntTacticDeclaration()

Returns

string

getSampler2DTacticDeclaration()

Returns

string

getPixels([flip])

Parameters

Name Type Description
flip Boolean Optional

Returns

Uint8ClampedArray Promise.<Uint8ClampedArray>

a Promise under the async contract, so await kernel.getPixels() is portable across every backend including webgpu, where the readback is genuinely async

updateTextureArgumentRefs(kernelValue, arg)

Parameters

Name Type Description
kernelValue WebGLKernelValue
arg GLTexture

Returns

Void

getType()

A ternary with an integer consequent but a float alternate promotes to float (the WGSL node's rule); the type system must agree with what exprConditional emits or the enclosing expression converts wrongly.

Returns

Void

toString()

FunctionBuilder drives tracing through toString(); for this backend that is the analysis pass — no text exists, the return value is always ''.

Returns

Void

emitFunction(assembler)

Bytecode pass. assembler carries the module builder, the kernel's baked memory layout and the shared global indices; it changes per size signature, so this may run repeatedly on one node.

Parameters

Name Type Description
assembler Object

Returns

Void

coerce()

Converts the wasm value on the stack top between categories. bool is an i32 constrained to 0/1, so bool→i32 is free and i32→bool renormalizes.

Returns

Void

emitByType()

Emits ast guaranteed to leave want on the stack, choosing the cast path by the type system's verdict, like the WGSL node's per-case castValue/castLiteral dispatches.

Returns

Void

emitCondition()

JS truthiness for conditions: comparisons pass through, numbers test against zero. Guards short-circuit code from ever seeing a raw f32.

Returns

Void

declareVecLocal()

Array(n) locals live as n consecutive f32 locals — wasm has no aggregate values outside memory, and these never escape the function.

Returns

Void

stmtRootReturn()

The root kernel stores into the output region at data_index (a shared global the run loop advances) and returns — the wasm-level return makes JS early returns exact at any nesting depth.

Returns

Void

forLoopIsSafe()

The WGSL node's exact safe-loop criteria: literal-init single declarator, safe test and init, both test and update present. Shared with the SIMD walk so the LOOP_MAX cap fires identically on both paths.

Returns

Void

stmtSwitch()

The switch lowers to an if/else chain on a discriminant local, exactly like the WGSL node: a case-terminating break is consumed, empty cases fall through by OR-ing their tests into the next case, a non-final default moves to the chain's end.

Returns

Void

collectSwitchGroups()

Fallthrough-empty-case grouping shared between the scalar and SIMD lowering: empty cases OR their tests into the next non-empty case, a non-final default moves to the chain's end.

Returns

Void

expression()

Emits ast, leaving exactly one value on the stack; returns its wasm category: 'f32' | 'i32' | 'bool' (i32 0/1) | 'void'.

Returns

Void

exprLogical()

Short-circuit is load-bearing, not an optimization: the right side may guard an out-of-range memory read (x > 0 && a[x - 1] > 0).

Returns

Void

emitMathCall()

Math.* calls compute in f32; native wasm opcodes where they exist, imports only for what the module actually uses (the kernel scans usedMathImports after analysis). All arguments go through the float ladder, matching the WGSL node's math-call casting.

Returns

Void

emitFlatLoad()

Flat row-major load, index = x + sizeX * (y + sizeY * z) — the same formula as the GL path's get32 and web-gpu's get_user_X, with missing y/z as zero. Dims and the region offset are baked; the kernel rebuilds per size signature.

Returns

Void

emitClampScalarIndex()

Clamps the i32 index on the stack top into [0, max]. i32 has no native min/max, so two selects.

Returns

Void

emitVectorFunction()

Bytecode pass for kernel_simd. Root kernel only; may run once per size signature like emitFunction.

Returns

Void

vAnalyze()

Emission-time variance analysis, run once per node and cached. Fixpoint over the local set: a local is varying when it is ever assigned a lane-varying value OR assigned at all under lane-varying control (divergent branch, varying-trip loop, ternary branch, short-circuit right side). A loop is varying-controlled when its test is varying or a break/continue reaches it from under a varying condition.

Returns

Void

vRecomputeCur()

Rebuilds vCur from a saved base after a construct. Retire masks are monotone accumulators within their scope, so base & ~each is exact at any later point; break/continue always target the innermost loop and the analysis forces any loop with a masked exit to be a varying loop, so only the top varying entry's masks apply.

Returns

Void

vSetLocal()

Stores the stack top into a v128 local; under a live branch mask the inactive lanes keep their previous value. bitselect copies exact bit patterns, so predication never perturbs IEEE results.

Returns

Void

vexprMask()

Leaves an i32x4 lane mask (all-ones/all-zeros) for ast as a condition; lane truth matches the scalar emitCondition exactly.

Returns

Void

vstatementBody()

A break/continue/return retires every lane that reached it, so the rest of its block is dead on both paths; the walk stops emitting there, and termination never leaks past the construct boundary.

Returns

Void

vStoreOutput()

Stores the quad's output. componentCount 1 is 4 consecutive f32 — one v128 store (load+blend+store when a mask is live). componentCount n is lane-strided, so components store scalar with a per-lane select.

Returns

Void

vexpr()

Emits ast in the vector walk. Uniform expressions go through the scalar walk untouched (one shared value; splatted only where a varying context needs it) and return scalar categories; varying expressions return 'vf32' | 'vi32' | 'vbool' (i32x4 lane mask) | 'void'.

Returns

Void

vexprShift()

i32x4 shifts take ONE scalar count for all lanes; a lane-varying count lane-scalarizes through the scalar opcode (same mod-32 masking).

Returns

Void

vexprLogical()

Predicated logic evaluates BOTH operands as masks (per-lane skipping cannot exist); side effects in the right operand still predicate correctly because vCur narrows to the left verdict while it runs, and the gather clamps keep formerly short-circuit-guarded reads from trapping.

Returns

Void

vUserCall()

Helpers stay scalar; a varying call lane-scalarizes: per lane, set that lane's thread.x and PCG state, extract the lane's arguments, call, and rebuild the result vector. Same function bodies and imports as the scalar path, so every lane is bit-identical to its scalar run.

Returns

Void

vGather()

Lane-varying gather: flat row-major index in i32x4 (the scalar formula lane-wise), clamped into the region so lanes a divergent branch turned off cannot trap, then 4 scalar loads + lane inserts — v128 has no gather. In-bounds lanes are untouched by the clamp.

Returns

Void

isThreadDependent()

Conservative thread-dependence for the SIMD phase: thread.x is the lane axis (thread.y/z are uniform across an x-row), Math.random is per-cell, user helper calls may read thread state internally, tainted locals propagate in walk order and never clear.

Returns

Void

new WebAssemblyKernel()

Returns

Void

WebAssemblyKernel.dispatchSpans()

SIMD quads must not cross an x-row (thread.y/z are uniform per quad): rows a multiple of 4 wide vectorize in one span, otherwise each row gets a vector span plus a scalar epilogue for its remainder cells. Shared with the pipeline executor, which drives per-step instances directly.

Returns

String

the path taken, for _lastRunPath

WebAssemblyKernel.translateSource()

The analysis pass: FunctionBuilder's trace runs each node's toString(), which for this backend resolves types and collects math-import/random usage without emitting a byte. Returns false for a return type this backend cannot store, so build() can degrade to cpu.

Returns

Void

WebAssemblyKernel.computeLayout()

[ args | constants | output ], each record 16-byte aligned, flat f32 (scalars are one 4-byte slot, Integer/Boolean viewed as i32). Dims come from the actual argument values, so the layout is per size signature.

Returns

Void

WebAssemblyKernel._threadable()

Threaded only under the async contract, only when threads exist, and only when the output is big enough (4096 cells) that splitting beats the postMessage round trip.

Returns

Void

WebAssemblyKernel._entryKey()

Sharedness is part of the cache key: a wasm memory import declares shared or not at compile time, so the same size signature needs a distinct module when the async contract routes it to the pool.

Returns

Void

WebAssemblyKernel._assembleModule()

The bytecode pass plus the run(start, end, seed) driver. The driver derives thread ids from the flat cell index with baked output dims (x fastest: x + sizeX * (y + sizeY * z), the storage order every backend shares) and seeds the PCG state per cell.

Returns

Void

WebAssemblyKernel._emitRunSimd()

run_simd(start, end, seed): 4 consecutive x cells per step. The caller guarantees (end - start) % 4 == 0 AND that no quad crosses an x-row, so thread.y/z are uniform per quad and thread.x is base + [0,1,2,3].

Returns

Void

WebAssemblyKernel._emitPcgRandomVector()

The scalar pcg_random lane-wise on an i32x4 state. Uniform-count shifts vectorize; the RXS shift count is per-lane, so that one step runs through the scalar opcodes per lane. The mask parameter predicates the state advance: a draw evaluated for an inactive lane (the untaken side of a divergent branch) must not advance that lane's stream, or every later draw in reconverged code desynchronizes from the scalar run.

Returns

Void

WebAssemblyKernel._emitPcgRandom()

PCG (permuted congruential, RXS-M-XS output) — the web-gpu kernel's pcg_random verbatim in i32 ops: bit-exact across platforms, top 24 bits scale into [0, 1) at full f32 mantissa resolution.

Returns

Void

WebAssemblyKernel._releaseEntry()

Frees everything an entry pins. The Memory itself has no explicit free, but dropping every reference (including the workers' — their instantiations hold the shared buffer) is the most a library can do to let it die young (#870). A shared entry defers until the threaded tail settles so an in-flight dispatch keeps what it captured.

Returns

Void

WebAssemblyKernel.checkArgumentTypes()

Returns

Void

WebAssemblyKernel._runThreaded()

The pool path, only ever reached under the async contract. Two timing constraints shape it: arguments must be sampled at CALL time (the sync contract's semantics), but the shared args region may still be feeding an in-flight run — so arguments flatten into staging copies now and are copied into wasm memory only when this run's turn on the memory comes up (_threadedTail). The seed is also drawn now so an unseeded kernel reseeds per call, not per settlement order.

The split follows the contract: min(pool, ceil(cells/4096)) contiguous ranges, each start aligned down to a multiple of 4 so every worker can enter run_simd; the last worker absorbs the tail.

Returns

Void

WebAssemblyKernel._shapeOutput()

The output region is tightly packed, so scalar returns reuse the memory-optimized erectors; Array(n) returns are stride n, shaped locally — the web-gpu kernel's exact conventions.

Returns

Void

VAL_TYPES()

Returns

Void

uleb()

Unsigned LEB128. Values are coerced through >>> so i32 bit patterns passed as negative JS numbers encode as their u32 counterpart.

Returns

Void

sleb()

Signed LEB128 for i32 immediates. The |0 coercion pins the value into i32 range so constants supplied as u32 bit patterns (0x9E3779B9-style hash multipliers) encode to the same 32 bits.

Returns

Void

uleb5At()

Fixed-width 5-byte ULEB128, written into an existing buffer. Used for call-target patch slots whose value is unknown when the body is emitted.

Returns

Void

utf8()

Minimal UTF-8 encoder so the builder stays dependency-free in both Node and the browser bundle (no Buffer, no assumed TextEncoder).

Returns

Void

new WasmFunctionEmitter()

Returns

Void

WasmFunctionEmitter.addLocal()

Returns

number

local index (params occupy the leading indices)

WasmFunctionEmitter.call()

Target is a NAME (import or defined function); the index is patched in at toBytes() so declaration order never matters.

Returns

Void

WasmFunctionEmitter.v128Const(lanes)

Parameters

Name Type Description
lanes ArrayLike.<number>

16 bytes, little-endian lane order

Returns

Void

new WasmModuleBuilder()

Returns

Void

WasmModuleBuilder.addMemoryImport()

One imported memory (env.memory) so a single compiled module can bind either a plain or a shared WebAssembly.Memory. Shared memories require a maximum by spec.

Returns

Void

WasmModuleBuilder.addFuncImport()

Returns

number

function index (imports lead the index space)

WasmModuleBuilder.addGlobal()

Returns

number

global index

WasmModuleBuilder.addFunction(name, signature)

Parameters

Name Type Description
name String
signature Object
signature.params Array.<String> Optional
signature.results Array.<String> Optional
signature.locals Array.<String> Optional

Returns

WasmFunctionEmitter

body emitter; more locals via addLocal()

WasmModuleBuilder.toBytes()

Returns

Uint8Array

the complete module binary

WORKER_SOURCE()

The worker body. A template string with NO closure captures: it must survive being evaluated from a blob URL (browser) or eval: true (Node), where nothing from this module's scope exists. Everything a task needs arrives by message: the compiled WebAssembly.Module and the shared WebAssembly.Memory structured-clone once per (worker, entry), then {start, end, seed} per task — results are written straight into the shared memory, so acks carry no data.

The run dispatch mirrors the kernel's sync path: quads must not cross an x-row, so run_simd is used only when the row width is a multiple of 4 (a 4-aligned range start then lands every quad inside one row); other shapes take the scalar export, which is bit-identical by the SIMD contract.

Pipeline entries ('pipelineSetup'/'pipelineRun') execute a WHOLE fused plan per task: every step module is instantiated over the plan's shared memory once at setup, then one run message walks all steps with an Atomics barrier between them — a generation counter in the shared memory, so step boundaries cost no postMessage round trip. Waits are sliced to 100ms so a barrier that can never fill (a peer died) is escapable: the main thread sets the abort word and notifies the generation word, and every check of either releases the worker to ack and go idle.

Returns

Void

new WebAssemblyWorkerPool()

Returns

Void

WebAssemblyWorkerPool._updateRef()

Event-loop handle accounting, per worker: ref'd while it has a setup OR a task in flight, unref'd when idle. The setup window matters -- a pending Promise does not hold Node's event loop, so a pool unref'd during module setup lets the process exit silently mid-dispatch. Browser workers have no ref/unref and need none.

Returns

Void

WebAssemblyWorkerPool._ensureSetup()

One setup message per (worker, entry) — concurrent tasks for the same entry share the in-flight ready wait rather than re-sending the module. Kernel entries and pipeline entries share this bookkeeping (a pool is owned by exactly one kernel or one pipeline executor, and pipeline ids are string-prefixed, so the id spaces cannot collide); only the setup message shape differs.

Returns

Void

WebAssemblyWorkerPool.dispatch(entry, tasks)

Parameters

Name Type Description
entry Object

kernel module entry: {id, module, memory, mathImports, sizeX}

tasks Array

contiguous {start, end, seed} ranges, one per worker

Returns

Promise.<undefined>

resolves when every range has been computed into the entry's shared memory

WebAssemblyWorkerPool.dispatchPipeline(entry, run)

One task per worker for a WHOLE fused plan: the barrier between steps lives in the entry's shared memory, so this is the only postMessage round trip a pipeline call makes. Every worker in [0, workerCount) must receive its task — the barrier fills only at workerCount arrivals — and a worker that dies rejects its task through the pool's usual machinery, which is the caller's signal to set the entry's abort word.

Parameters

Name Type Description
entry Object

pipeline entry: {id, pipeline, memory, modules, moduleMathImports, steps, countIndex, genIndex, abortIndex, workerCount, workerRanges}

run Object

per-call inputs: {baseGen, seeds}

Returns

Promise.<undefined>

resolves when every worker has acked its walk of the plan

WebAssemblyWorkerPool.release()

Drops an entry's instantiation from every live worker: the worker-side instances are what keep an evicted entry's shared memory alive (#870). The caller guarantees no task for this entry is still in flight.

Returns

Void

SUPPORTED_VALUE_TYPES()

Fused pipeline execution (docs/design/pipeline-compilation.md): every plan step compiles to a wasm module over ONE shared memory laid out [ barrier control | pipeline args | literals | constants | plan buffers ] (the control words exist only on the threaded path), with each module's input/output offsets baked against that layout. Per call there is one flattenTo per pipeline argument and one readback per result, however many steps the plan unrolls to.

Sync path ('fused-sync'): steps run back-to-back on the calling thread.

Threaded path ('fused-threaded'): when wasm threads exist and the plan is big enough, pool workers execute the WHOLE plan — each worker owns a contiguous cell-range slice of every step and advances step-to-step on an Atomics barrier (generation counter in the shared memory). The main thread dispatches once per call and then waits only for the final generation, so step boundaries cost no main-thread round trip. Threads unavailable or the plan too small falls back to the sync path; anything the backend cannot take at all degrades to the generic executor as usual.

Returns

Void

new FusionFallback()

The degradation signal, per the backend's usual contract: the pipeline catches it and runs the generic executor with this reason. recompilable marks argument size/type drift a fresh fused compile can absorb.

Returns

Void

FusionFallback.compile(pipeline, plan, args)

Parameters

Name Type Description
pipeline Pipeline
plan Object
  • phase-1 plan IR; buffer assignment is reused as-is
args Array
  • the first call's sampled arguments; their sizes and types bake into the layout, and execute() re-checks them per call

Returns

WebAssemblyPipelineExecutor

this._extraShortcuts()

kernels created for second and later type signatures of one plan kernel (the plan clone carries the first); destroyed with the executor

Returns

Void

_compile()

A program is a plan kernel prepared for one argument-type signature: type inference and bytecode translation ran, but no module was instantiated — modules are per step-offset assignment, built below.

Returns

Void

_representativeArgs()

Stand-ins with the exact types and dims each binding will have at run time, for setupArguments/computeLayout: sampled values stand for themselves, a step output becomes an Input over its producer's dims.

Returns

Void

_prepareKernel()

The analysis half of WebAssemblyKernel.build() without instantiation: modules are assembled against the shared layout instead. Inference is reset first — on a fused recompile the same kernel must re-infer for the new signature, not keep the old one.

Returns

Void

_checkArguments()

The layout baked argument sizes and scalar types; a call that drifts from them throws recompilable so the pipeline compiles a fresh fused plan for the new signature, the way the kernel itself re-instantiates per size signature.

Returns

Void

execute(args)

Parameters

Name Type Description
args Array
  • sampled pipeline arguments

Returns

results shaped per the plan; synchronous on the sync path (the pipeline's tail promise provides the async contract), a Promise on the threaded path

_executeThreaded()

One pool dispatch for the whole plan; the workers walk every step over the already-written args and meet at the memory-resident barrier, so the only thing left to await here is the final generation. The pipeline tail serializes calls, which is what makes resetting the generation counter safe: no worker touches the control words between a run's final barrier and its next run message.

Returns

Void

_waitForGeneration()

Resolves when the generation counter reaches target, rejects on abort or when the counter stalls past sanityTimeoutMs. Atomics.waitAsync where the host has it (woken by the workers' notify and by _abort), short-slice polling otherwise — either way the main thread never blocks.

Returns

Void

_abort()

Releases every wait on the run: workers poll the abort word at each barrier (and inside their sliced Atomics.wait), the main thread checks it on every generation wake. First cause wins; the executor is dead afterwards — the barrier count is indeterminate.

Returns

Void

abortRuns()

Entry point for Pipeline.destroy() while a run may be in flight: the sync path cannot be mid-run (it never yields), so only the threaded path has anything to interrupt.

Returns

Void

_readResults()

The one readback: slice copies results out of wasm memory only here. On the threaded path the barrier's final generation happened-before this read, so the workers' stores are visible.

Returns

Void

astFunction(ast, retArr)

Parameters

Name Type Description
ast Object
  • the AST object to parse
retArr Array
  • return array string

Returns

Array

the append retArr

astReturnStatement(ast, retArr)

Parameters

Name Type Description
ast Object
  • the AST object to parse
retArr Array
  • return array string

Returns

Array

the append retArr

astLiteral(ast, retArr)

Parameters

Name Type Description
ast Object
  • the AST object to parse
retArr Array
  • return array string

Returns

Array

the append retArr

astBinaryExpression(ast, retArr)

Parameters

Name Type Description
ast Object
  • the AST object to parse
retArr Array
  • return array string

Returns

Array

the append retArr

castLiteralToInteger(ast, retArr)

Parameters

Name Type Description
ast Object
retArr Array

Returns

Array.<String>

castLiteralToFloat(ast, retArr)

Parameters

Name Type Description
ast Object
retArr Array

Returns

Array.<String>

castValueToInteger(ast, retArr)

Parameters

Name Type Description
ast Object
retArr Array

Returns

Array.<String>

castValueToFloat(ast, retArr)

Parameters

Name Type Description
ast Object
retArr Array

Returns

Array.<String>

astIdentifierExpression(idtNode, retArr)

Parameters

Name Type Description
idtNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the append retArr

markupUserName()

Returns

Void

astForStatement(forNode, retArr)

Parameters

Name Type Description
forNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the parsed webgl string

loopIndexAssignedInLoop()

Returns

Void

astWhileStatement(whileNode, retArr)

Parameters

Name Type Description
whileNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the parsed webgl string

astDoWhileStatement(doWhileNode, retArr)

Parameters

Name Type Description
doWhileNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the parsed webgl string

astAssignmentExpression(assNode, retArr)

Parameters

Name Type Description
assNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the append retArr

astBlockStatement(bNode, retArr)

Parameters

Name Type Description
bNode Object
  • the AST object to parse
retArr Array
  • return array string

Returns

Array

the append retArr

traceFunctionAST(ast)

Parameters

Name Type Description
ast Object
  • the parsed function node

Returns

Void

normalizeBlock(block)

Parameters

Name Type Description
block Object
  • a BlockStatement whose body may be rewritten

Returns

Void

normalizeBranch(statement, key)

Parameters

Name Type Description
statement Object
  • the owning statement
key String
  • which branch to normalize, braced first if needed

Returns

Void

normalizeLoopHeader(statement)

Parameters

Name Type Description
statement Object
  • the loop node

Returns

Object

a replacement BlockStatement, or null when the header is clean and the loop should be left as written

stampSyntheticNodes()

Returns

Void

astVariableDeclaration(varDecNode, retArr)

Parameters

Name Type Description
varDecNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the append retArr

astIfStatement(ifNode, retArr)

Parameters

Name Type Description
ifNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the append retArr

astSwitchCaseConsequent(consequent, retArr)

Parameters

Name Type Description
consequent Array
  • the case's statements
retArr Array
  • return array string

Returns

Array

the append retArr

astThisExpression(tNode, retArr)

Parameters

Name Type Description
tNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the append retArr

astMemberExpression(mNode, retArr)

Parameters

Name Type Description
mNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the append retArr

astCallExpression(ast, retArr)

Parameters

Name Type Description
ast Object
  • the AST object to parse
retArr Array
  • return array string

Returns

Array

the append retArr

astArrayExpression(arrNode, retArr)

Parameters

Name Type Description
arrNode Object
  • the AST object to parse
retArr Array
  • return array string

Returns

Array

the append retArr

nodeIsSideEffectFree(node)

Parameters

Name Type Description
node Object
  • the expression node

Returns

Boolean

statementContainsNestedIndexRead(statement)

Parameters

Name Type Description
statement Object
  • the statement node

Returns

Boolean

containsCallTo(node, name)

Parameters

Name Type Description
node Object
  • the node to search
name String
  • the callee name

Returns

Boolean

containsNestedSameFunctionCall(statement)

Parameters

Name Type Description
statement Object
  • the node to search

Returns

Boolean

statementIsSideEffectFreeBesidesTopLevelAssignment(statement)

Parameters

Name Type Description
statement Object
  • the statement node

Returns

Boolean

testCanvas()

Returns

Void

testContext()

Returns

Void

new WebGLKernel()

Properties

Name Type Description
textureCache Array.<WebGLTexture>
  • webGl Texture cache
programUniformLocationCache Object.<string, WebGLUniformLocation>
  • Location of program variables in memory
framebuffer WebGLFramebuffer
  • Webgl frameBuffer
buffer WebGLBuffer
  • WebGL buffer
program WebGLProgram
  • The webGl Program
functionBuilder FunctionBuilder
  • Function Builder instance bound to this Kernel
pipeline Boolean
  • Set output type to FAST mode (GPU to GPU via Textures), instead of float
endianness string
  • Endian information like Little-endian, Big-endian.
argumentTypes Array.<string>
  • Types of parameters sent to the Kernel
compiledFragmentShader string
  • Compiled fragment shader string
compiledVertexShader string
  • Compiled Vertical shader string

Returns

Void

WebGLKernel.lookupKernelValueType(type, dynamic, precision, value)

Parameters

Name Type Description
type
dynamic
precision
value

Returns

KernelValue

WebGLKernel.constructor(source, settings)

Parameters

Name Type Description
source String IKernelJSON
settings IDirectKernelSettings

Returns

Void

this.maxTexSize()

Returns

Void

this.threadDim()

The thread dimensions, x, y and z

Returns

Void

initContext()

Returns

WebGLRenderingContext

initPlugins(settings)

Parameters

Name Type Description
settings IDirectKernelSettings

Returns

Array.<string>

validateSettings(args)

Parameters

Name Type Description
args IArguments

Returns

Void

deleteTexture(texture)

Parameters

Name Type Description
texture WebGLTexture

Returns

Void

_replaceOutputTexture()

Returns

Void

_setupOutputTexture()

Returns

Void

_replaceSubOutputTextures()

Returns

Void

_setupSubOutputTextures()

Returns

Void

getUniformLocation()

Returns

Void

_getFragShaderArtifactMap(args)

Parameters

Name Type Description
args Array
  • The actual parameters sent to the Kernel

Returns

Object

An object containing the Shader Artifacts(CONSTANTS, HEADER, KERNEL, etc.)

_getVertShaderArtifactMap(args)

Parameters

Name Type Description
args Array
  • The actual parameters sent to the Kernel

Returns

Object

An object containing the Shader Artifacts(CONSTANTS, HEADER, KERNEL, etc.)

_getHeaderString()

Returns

String

result

_getLoopMaxString()

Returns

String

result

_getConstantsString()

Returns

String

result

_getTextureCoordinate()

Returns

String

result

_getDecode32EndiannessString()

Returns

String

result

_getEncode32EndiannessString()

Returns

String

result

_getDivideWithIntegerCheckString()

Returns

String

result

_getMainArgumentsString(args)

Parameters

Name Type Description
args Array
  • The actual parameters sent to the Kernel

Returns

String

result

getKernelString()

Returns

String

result

getMainResultKernelPackedPixels()

Returns

String

getMainResultSubKernelPackedPixels()

Returns

String

replaceArtifacts(src, map)

Parameters

Name Type Description
src String
  • Shader string
map Object
  • Variables/Constants associated with shader

Returns

Void

getFragmentShader(args)

Parameters

Name Type Description
args Array
  • The actual parameters sent to the Kernel

Returns

string

Fragment Shader string

getVertexShader(args)

Parameters

Name Type Description
args Array IArguments
  • The actual parameters sent to the Kernel

Returns

string

Vertical Shader string

toString()

Returns

Void

toJSON()

Returns

IKernelJSON

new WebGL2FunctionNode()

Returns

the converted webGL function string

WebGL2FunctionNode.astIdentifierExpression(idtNode, retArr)

Parameters

Name Type Description
idtNode Object
  • An ast Node
retArr Array
  • return array string

Returns

Array

the append retArr

testCanvas()

Returns

Void

testContext()

Returns

Void

features()

Returns

Void

new WebGL2Kernel()

Returns

Void

WebGL2Kernel.getFeatures()

Returns

IKernelFeatures

initContext()

Returns

WebGLRenderingContext WebGL2RenderingContext

validateSettings(args)

Parameters

Name Type Description
args IArguments

Returns

Void

renderValues()

Returns

Void

renderOutputAsync()

Returns

Void

_getHeaderString()

Returns

String

result

_getTextureCoordinate()

Returns

String

result

_getMainArgumentsString(args)

Parameters

Name Type Description
args Array
  • The actual parameters sent to the Kernel

Returns

String

result

getKernelString()

Returns

String

result

getMainResultKernelPackedPixels()

Returns

String

getMainResultSubKernelPackedPixels()

Returns

String

toJSON()

Returns

IKernelJSON

new WebGPUBufferResult()

Returns

Void

WebGPUBufferResult.constructor(settings)

Parameters

Name Type Description
settings Object
settings.buffer GPUBuffer
settings.output Array.<number>

logical dims, e.g. [512, 512]

settings.componentCount number

1 for scalar returns, 2/3/4 for Array(n)

Optional
settings.context Object

the {adapter, device} pair from WebGPUContext

settings.kernel Object

owning WebGPUKernel, used for readback + shaping

Returns

Void

WebGPUBufferResult.toArray()

Returns

Promise.<Float32Array|Array>

values shaped exactly as a non-pipeline run of the producing kernel would resolve them

contextPromise()

Returns

Void

acquire()

Returns

Promise.<{ adapter: GPUAdapter, device: GPUDevice }>

cached and shared

destroy()

Destroys the shared device. gpu.destroy() does NOT call this — the device outlives any one GPU instance; this is the page-level teardown.

Returns

Promise.<undefined>

new WGSLFunctionNode()

Returns

Void

WGSLFunctionNode.wgslFloat(value)

WGSL rejects float literals that overflow f32 (GLSL forgave them), and JS toString of an integral double has no decimal point. Every float literal goes through here so the emitted text is always a committed, in-range WGSL float.

Parameters

Name Type Description
value number

Returns

String

WGSLFunctionNode.wgslInt(value)

Parameters

Name Type Description
value number

Returns

String

WGSLFunctionNode.mangleFunctionName(name)

User function names collide with WGSL keywords and builtins where GLSL names did not; mangle rather than reject. Unconditionally: WGSL reserves over sixty words that are legal JavaScript function names (filter, get, set, type, self, ...), and a curated list drifts out of date with the spec -- #861 found 64 missing. The fn_ prefix removes the class the way user_ already does for variables; the registry side (FunctionBuilder, type inference) keys on original names and never sees this.

Parameters

Name Type Description
name String

Returns

String

WGSLFunctionNode.getLookupType()

A WebGPUBufferResult argument reads like a flat storage array; the base typeLookupMap does not know the type, so value lookup is resolved here.

Returns

Void

WGSLFunctionNode.astUpdateExpression()

WGSL has no ternary. Pure-value case lowers to select(false, true, cond) — both sides evaluate eagerly, which is observationally safe in the side-effect-free kernel language. Void (minified) case lowers to if/else.

Returns

Void

WGSLFunctionNode.astFunction()

Root kernel: emits only body statements — the kernel class assembles the @compute entry, guard and data_index around them. Non-root: a full fn name(args) -> type { ... } declaration.

Returns

Void

WGSLFunctionNode.checkAndUpconvertOperator()

** upconverts to pow(f32, f32); bitwise operators are native in WGSL (GLSL ES 1.00 needed helper functions) and need only operand casts.

Returns

Void

this._extraShortcuts()

kernels created for second and later type signatures of one plan kernel (the plan clone carries the first); destroyed with the executor

Returns

Void

_representativeArgs()

Stand-ins with the exact types and dims each binding will have at run time, for the kernel build: sampled values stand for themselves, a step output becomes an Input over its producer's dims.

Returns

Void

_checkArguments()

The layout baked argument sizes and scalar types; a call that drifts from them throws recompilable so the pipeline compiles a fresh fused plan for the new signature, the way the kernel itself rebuilds per size signature.

Returns

Void

execute(args)

Parameters

Name Type Description
args Array
  • sampled pipeline arguments

Returns

Promise.<*>

results shaped per the plan. Argument drift throws a recompilable FusionFallback synchronously, before anything is encoded; async failures (device loss) reject, which drops the executor via the pipeline's _guardAsync so the next call compiles fresh.

new WebGPUKernel()

Returns

Void

destroyContext()

The shared device is a module singleton that outlives any one GPU instance; page-level teardown goes through WebGPUContext.destroy().

Returns

Void

build()

Everything device-independent — validation, JS→WGSL translation, module assembly — happens synchronously here, so unsupported constructs throw at the first call rather than rejecting; the device round trip continues in _buildAsync and run() chains on its promise.

Returns

Void

computeParamsLayout()

Params struct, matching the WGSL struct member for member: outputX/outputY/outputZ/_pad0, one vec4 [sizeX, sizeY, sizeZ, total] per array argument, then scalar arguments as f32/i32/u32, the whole buffer padded to a 16-byte multiple.

Returns

Void

_computeDispatch(threadDim)

Parameters

Name Type Description
threadDim Array.<number>

Returns

[object Object]

_checkBufferSize()

Oversized bindings must throw here: past the device limit, bind-group validation fails asynchronously, the submit is dropped, and the zero- initialized staging buffer would resolve a fully-shaped all-zeros result — silent wrong data instead of an error.

Returns

Void

_snapshotArguments()

The mutation-sensitive step: array arguments are flattened into fresh Float32Arrays synchronously inside run(), before any await, preserving the GL path's snapshot semantics.

Returns

Void

_runInternal()

Steady state is synchronous through queue.submit — writeBuffer copies its data at call time and one queue executes submits in order, so overlapping un-awaited calls cannot race; only the readback awaits, on a staging buffer of its own.

Returns

Void

_acquireStaging()

A staging buffer is unusable from mapAsync until unmap, so overlapping un-awaited calls each need their own; sequential awaited calls reuse one buffer forever. Capped so a burst cannot ratchet memory.

Returns

Void

_shapeOutput()

Readback is tightly packed, so scalar returns reuse the memory-optimized erectors; Array(n) returns are stride n (not the GL texel stride 4), shaped locally.

Returns

Void

readBufferResult(handle)

Readback for a pipeline handle: copy → map → shape, same conventions as a non-pipeline run of the producing kernel.

Parameters

Name Type Description
handle WebGPUBufferResult

Returns

Promise.<Float32Array|Array>

getPixels([flip])

Parameters

Name Type Description
flip Boolean Optional

Returns

Promise.<Uint8ClampedArray>

new GLTexture()

Properties

Name Type Description
framebuffer

Returns

Void

GLTexture.textureType()

Returns

Number

GLTexture.beforeMutate()

Returns

Boolean

GLTexture.cloneTexture() private method

Returns

Void

GLTexture.newTexture() private method

Returns

Void

new WebGLKernelArray()

Returns

Void

WebGLKernelArray.rebind()

Puts this value's texture back on its unit. Texture units are context state shared by every kernel, and each kernel numbers its own from zero, so between runs another kernel's textures sit on them (#862). The data in this texture is intact -- only the binding needs to come back.

Returns

Void

WebGLKernelArray.getBitRatio(value)

bit storage ratio of source to target 'buffer', i.e. if 8bit array -> 32bit tex = 4

Parameters

Name Type Description
value

Returns

number

constructor(value, settings)

Parameters

Name Type Description
value KernelVariable
settings IWebGLKernelValueSettings

Returns

Void

rebind()

Re-establishes whatever this value put on shared context state before a run. Scalar uniforms live on the program and need nothing.

Returns

Void

getStringValueHandler()

Used for when we want a string output of our kernel, so we can still input values to the kernel

Returns

Void

updateValue(inputTexture)

Parameters

Name Type Description
inputTexture GLTextureMemoryOptimized

Returns

Void

updateValue(inputTexture)

Parameters

Name Type Description
inputTexture GLTexture

Returns

Void