Components
Every component in the Studio palette, grouped by category, with its fields and defaults.
Components This page catalogs every component on the Layers tab, grouped the way the sidebar groups them, with each field and its default value exactly as Studio shows them. The Input, Dataset, and Output nodes have their own page: see Inputs and outputs /docs/studio/inputs-and-outputs . Datasets live on the Data tab and blocks on the Blocks tab; see Blocks /docs/studio/blocks . A layer connected straight to the Output node has its Activation field locked to Linear, with the note "Final activation is controlled by the Output node."; its earlier value comes back when you disconnect it. This applies to Dense Layer , Conv Layer , Separable Conv , MoE Layer , and FFN Block . Layers Label What it does Fields and defaults ------------------- ------------------------------------------------------------------------------------------------------------------------------------ --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- Dense Layer Fully connected layer. Requires 1D input use Global Pooling or Flatten first . Units 128, Activation ReLU ReLU, Sigmoid, Tanh, Softmax, Linear , Use Bias on Conv Layer Applies convolutional filters to extract spatial features. Dimensions 2d 1d, 2d, 3d , Number of Filters 32, Kernel Size 3, Stride 1, Padding Same Same, Valid, Explicit , Activation ReLU None, ReLU, Sigmoid, Tanh Transposed Conv Upsamples feature maps to increase spatial resolution. Used in decoder networks. Dimensions 2d 2d, 3d , Number of Filters 32, Kernel Size 3, Stride 2, Padding Same Same, Valid, Explicit Separable Conv Depthwise separable convolution. Far fewer parameters than a standard Conv layer. Dimensions 2d 1d, 2d , Output Filters 64, Kernel Size 3, Stride 1, Padding Same Same, Valid, Explicit , Depth Multiplier 1, Activation ReLU None, ReLU, ReLU6, Sigmoid, Tanh, Swish Upsample Upsamples spatial dimensions by interpolation. No learned parameters. Scale Factor 2, Interpolation Mode Nearest Nearest, Bilinear, Bicubic, Trilinear , Dimensions 2d 1d, 2d, 3d , Align Corners off VAE Reparam Reparameterization trick for variational autoencoders: samples a latent from a mean and variance. Latent Dimension 20 Noise Scheduler Defines the noise schedule for a diffusion model; outputs noisy data and the noise. Schedule Type Linear Linear, Cosine , Number of Timesteps 1000, Beta Start 0.0001, Beta End 0.02 CFG Blends conditional and unconditional outputs for guided diffusion generation. Guidance Scale 7.5, Unconditional Probability 0.1 MoE Layer Mixture-of-experts layer: parallel feed-forward networks with top-k routing for sparse transformers. Number of Experts 4, Model Dimension 256, FFN Dimension 1024, Top-K Experts 2, Activation SwiGLU ReLU, SwiGLU, GELU , Load Balance Weight 0.01 Router Expert routing for MoE architectures, by top-k softmax or noisy top-k selection. Number of Experts 4, Routing Type Top-K Softmax Top-K Softmax, Noisy Top-K , Capacity Factor 1.25 Mel Spec Converts a raw audio waveform into a mel-frequency spectrogram. FFT Size 1024, Hop Length 512, Mel Bands 64 MFCC Extracts mel-frequency cepstral coefficients from an audio waveform. Number of MFCCs 13, FFT Size 1024, Hop Length 512 LoRA Low-rank adaptation: parameter-efficient fine-tuning that adds a small trainable update to a layer. Rank 8, Alpha 16, Target Modules query, value query, key, value, output, dense FFN Block Dim-preserving transformer feed-forward block: expands to a wider hidden dimension, applies a gated activation, then contracts back. Model Dimension 256, FFN Dimension 1024, Activation SwiGLU ReLU, SwiGLU, GEGLU , Dropout 0 Embeddings & Encoding Label What it does Fields and defaults ----------------------- -------------------------------------------------------------------------------------------------- -------------------------------------------------------------------------------------------- Embedding Layer Converts integer token indices to dense vectors. Connect it to a Dataset node or a sequence Input. Vocabulary Size 10000, Embedding Dimension 256, Mask Zero off, Input Length maxlen 200 Positional Encoding Adds positional information to sequence input, fixed sinusoidal or trainable learned . Encoding Type Sinusoidal Sinusoidal, Learned , Max Sequence Length 512, Model Dimension 256 RoPE Rotary position embedding: relative positional encoding used in modern language models. Dimension 64, Base Frequency 10000, Max Sequence Length 2048 Patch Embedding Splits an image into patches and projects them to embeddings, for Vision Transformers. Patch Size 16, Embedding Dimension 768, Image Size 224 CLS Token Prepends a learnable classification token to the input sequence, for ViT-style models. Embedding Dimension 768 Timestep Embedding Embeds the diffusion timestep supplied by the training loop. Time Dimension 128, Max Period 10000, Hidden Dimension 512 Recurrent Label What it does Fields and defaults -------- ----------------------------------------------------------------------------------------------------- ----------------------------------------------------------------------------------------------------------- LSTM Long short-term memory layer for sequences. Turn off Return Sequences to output 1D for a Dense layer. Hidden Units 128, Number of Layers 1, Return Sequences on, Bidirectional off, Dropout 0, Use Projection off GRU Gated recurrent unit layer for sequences. Turn off Return Sequences to output 1D for a Dense layer. Hidden Units 128, Number of Layers 1, Return Sequences on, Dropout 0, Use Projection off Attention Label What it does Fields and defaults ------------------------ ---------------------------------------------------------------------------------------------------- ------------------------------------------------------------------------------------------------------------------ Multi-Head Attention Multi-head self-attention. Requires sequence input, for example from an Embedding or LSTM/GRU layer. Number of Heads 8, Key Dimension 64, Dropout Rate 0.1, Use Causal Mask off, KV Heads GQA/MQA 8, Use KV Cache off Cross-Attention A query sequence attends to a separate context sequence. Used in decoders and U-Net conditioning. Query Dimension 256, Context Dimension 256, Number of Heads 8, Dropout Rate 0.1 Activation Label What it does Fields and defaults -------------- ----------------------------------------------------------------- ------------------------------------------------------------------------------------------------------------------------------------------------------------------- Activation Applies a non-linear activation function element-wise. Activation Function ReLU ReLU, GELU, Sigmoid, Tanh, Softmax, Leaky ReLU, Clipped ReLU, ELU, Swish , Clip Value for Clipped ReLU 6, Alpha for LeakyReLU/ELU 0.3 GLU Gated linear unit. SwiGLU gates with SiLU; GeGLU gates with GELU. Variant SwiGLU SwiGLU, GeGLU Normalization Label What it does Fields and defaults ----------------- ------------------------------------------------------------------------------------ ----------------------------------------------------------------------- Layer Norm Normalizes inputs across the feature dimension. Works with any tensor type. Epsilon 1e-6, Center use beta on, Scale use gamma on Batch Norm Normalizes inputs across the batch dimension. Works with any tensor type. Momentum 0.99, Epsilon 1e-3, Center use beta on, Scale use gamma on Instance Norm Normalizes each sample independently. Used for style transfer and generative models. Epsilon 1e-5, Learnable Affine on RMS Norm Root mean square normalization, faster than Layer Norm. Epsilon 1e-6 Group Norm Normalizes inputs across channel groups. Works with any tensor type. Number of Groups 32, Epsilon 1e-5 Regularization Label What it does Fields and defaults -------------------- ---------------------------------------------------------------------------------------------- ---------------------------------------------------------------------------------------------- Dropout Randomly zeroes elements during training. Use a Spatial mode for CNNs to drop entire channels. Dropout Rate 0.5, Mode Standard Standard, Spatial 1D, Spatial 2D, Spatial 3D , Random Seed 42 Stochastic Depth Randomly drops entire blocks during training; acts as identity during evaluation. Drop Probability 0.1 L1 Applies an L1 weight penalty to promote weight sparsity. L1 Factor 0.01 L2 Applies an L2 weight penalty weight decay to reduce overfitting. L2 Factor 0.01 Pooling Label What it does Fields and defaults ------------- -------------------------------------------------------------------------------- ------------------------------------------------------------------------------------------------------------------------------------------- Pooling Downsamples spatial data. Use a Global option to flatten for a Dense layer. Pooling Type Max Max, Average, Global Max, Global Average , Dimensions 2d 1d, 2d, 3d , Pool Size 2, Stride 2, Padding valid valid, same Unpooling Reverses a pooling operation to upsample feature maps. Used in decoder networks. Dimensions 2d, Pool Size 2 Utility Label What it does Fields and defaults --------------------- -------------------------------------------------------------------------------------------------- ------------------------------------------------------------------------------------- Flatten Flattens multi-dimensional input into a 1D tensor, for example after conv layers and before Dense. Flatten From Dimension 0 Reshape Reshapes a tensor to a target shape. The total element count must stay the same. Target Shape "-1,256" Permute Reorders tensor dimensions, for example to convert channel-first to channel-last. Dimension Order "0,2,1" Merge Combines multiple inputs. Use Add for residual or skip connections. Operation Add Add, Concatenate, Multiply, Average , Concat Axis -1 for Concatenate Quantization Fake quantization during training, real quantization at inference, for deployment. Bit Width 8 4, 8 , Quantization Type Symmetric Symmetric, Asymmetric Extract CLS Token Extracts a single sequence position default 0, the CLS token for a classification head. Sequence Position 0
Open in Dagnam.AI docs