Skip to content

14. Keras Sequential and Functional APIΒΆ

Learn how to design Deep Learning models using Keras's Sequential and Functional APIs, understand when each approach should be used, build complex computation graphs, work with multiple inputs and outputs, reuse layers, implement skip connections, and design production-ready neural network architectures.


🎯 Learning Objectives¢

After completing this chapter, you will be able to:

  • Understand the Keras Sequential API
  • Understand the Keras Functional API
  • Compare Sequential and Functional model construction
  • Build simple neural networks using Sequential
  • Build complex neural networks using the Functional API
  • Understand Keras symbolic tensors
  • Understand the relationship between inputs, layers, and outputs
  • Build models with multiple inputs
  • Build models with multiple outputs
  • Reuse layers across different parts of a model
  • Build branching architectures
  • Build merging architectures
  • Implement skip connections
  • Understand residual-style connections
  • Inspect Functional API model graphs
  • Visualize model architecture
  • Share layers between models
  • Build reusable model components
  • Understand model composition
  • Understand when to choose Sequential, Functional API, or subclassing
  • Design maintainable Keras architectures for enterprise Deep Learning systems

πŸ“– OverviewΒΆ

Keras provides multiple ways to define neural network models.

The three major approaches are:

Sequential API
      β”‚
      β”œβ”€β”€ Simple linear stacks
      β”‚
Functional API
      β”‚
      β”œβ”€β”€ Complex computation graphs
      β”‚
Model Subclassing
      β”‚
      └── Maximum customization

The first two approaches are especially important for most Deep Learning applications.

flowchart TD

    KERAS["Keras Model Construction"]

    KERAS --> SEQ["Sequential API"]
    KERAS --> FUNC["Functional API"]
    KERAS --> SUB["Model Subclassing"]

    SEQ --> SIMPLE["Simple Layer Stack"]

    FUNC --> COMPLEX["Complex Graphs"]

    SUB --> CUSTOM["Custom Architecture / Behavior"]

🧠 Why Multiple Model APIs?¢

Different neural networks have different architectural requirements.

A simple classifier may look like:

Input
  ↓
Dense
  ↓
Dense
  ↓
Output

A more complex architecture may look like:

             β”Œβ”€β”€ Dense ──┐
Input ────────            β”œβ”€β”€ Merge ── Output
             └── Dense β”€β”€β”€β”˜

A residual network may contain:

Input ────────────────┐
  ↓                   β”‚
Layer                 β”‚
  ↓                   β”‚
Layer                 β”‚
  ↓                   β”‚
  └────── Add β—„β”€β”€β”€β”€β”€β”€β”€β”˜

The Sequential API is ideal for the first case.

The Functional API is designed for the second and third cases.


πŸ— Sequential APIΒΆ

The Sequential API represents a simple linear stack of layers.

flowchart LR

    INPUT["Input"]

    L1["Dense"]

    L2["Dense"]

    L3["Dense"]

    OUTPUT["Output"]

    INPUT --> L1
    L1 --> L2
    L2 --> L3
    L3 --> OUTPUT

Each layer receives the output of the previous layer.


πŸ§ͺ Basic Sequential ModelΒΆ

import tensorflow as tf


model = tf.keras.Sequential([

    tf.keras.Input(
        shape=(784,)
    ),

    tf.keras.layers.Dense(
        128,
        activation="relu"
    ),

    tf.keras.layers.Dense(
        64,
        activation="relu"
    ),

    tf.keras.layers.Dense(
        10,
        activation="softmax"
    )
])

The architecture is:

784 Features
     ↓
Dense(128)
     ↓
Dense(64)
     ↓
Dense(10)
     ↓
Softmax

🧠 Sequential Model Concept¢

The Sequential API can be viewed as:

[ f(x) = f_n( f_{n-1}( ... f_2( f_1(x) ) ... ) ) ]

Each layer transforms the output of the previous layer.


🧱 Adding Layers¢

Layers can also be added incrementally.

model = tf.keras.Sequential()

model.add(
    tf.keras.Input(
        shape=(784,)
    )
)

model.add(
    tf.keras.layers.Dense(
        128,
        activation="relu"
    )
)

model.add(
    tf.keras.layers.Dense(
        10,
        activation="softmax"
    )
)

🧠 When Sequential Works Well¢

Use Sequential when:

  • The architecture is linear
  • There is one input
  • There is one output
  • Every layer has exactly one input
  • Every layer produces exactly one output
  • There are no branching paths
  • There are no skip connections

Typical examples:

Simple Classification
Simple Regression
Basic MLP
Simple CNN
Basic Feed-Forward Network

⚠ Sequential API Limitations¢

Sequential becomes unsuitable when the architecture contains:

Multiple Inputs
Multiple Outputs
Branching
Layer Sharing
Skip Connections
Non-linear Graphs

For these architectures, use the Functional API.


🧠 Functional API¢

The Functional API represents a neural network as a directed computation graph.

Instead of saying:

model.add(layer)

we explicitly connect tensors:

x = layer(inputs)

and finally create:

model = tf.keras.Model(
    inputs=inputs,
    outputs=outputs
)

πŸ— Functional API ArchitectureΒΆ

flowchart LR

    INPUT["Input Tensor"]

    D1["Dense 128"]

    D2["Dense 64"]

    OUT["Output"]

    INPUT --> D1
    D1 --> D2
    D2 --> OUT

The important difference is that the engineer explicitly defines the graph.


πŸ§ͺ Basic Functional API ModelΒΆ

import tensorflow as tf


inputs = tf.keras.Input(
    shape=(784,)
)

x = tf.keras.layers.Dense(
    128,
    activation="relu"
)(inputs)

x = tf.keras.layers.Dense(
    64,
    activation="relu"
)(x)

outputs = tf.keras.layers.Dense(
    10,
    activation="softmax"
)(x)

model = tf.keras.Model(
    inputs=inputs,
    outputs=outputs
)

🧠 Functional API Mental Model¢

Think of:

inputs

as the starting node.

Each layer:

layer(x)

creates a new tensor.

Finally:

outputs

becomes the endpoint.

flowchart LR

    INPUT["inputs"]

    L1["Layer 1"]

    T1["Tensor"]

    L2["Layer 2"]

    T2["Tensor"]

    OUT["outputs"]

    INPUT --> L1
    L1 --> T1
    T1 --> L2
    L2 --> T2
    T2 --> OUT

🧠 Symbolic Tensors¢

When using the Functional API:

inputs = tf.keras.Input(
    shape=(784,)
)

inputs represents a symbolic tensor describing the expected computation.

It is not an actual batch of training data.

This distinction is important.

Functional API

Symbolic Tensor
      ↓
Describes Computation
      ↓
Keras Builds Graph

During training:

Actual Data
      ↓
Computation Graph
      ↓
Output

🧠 Sequential vs Functional¢

Feature Sequential Functional
Simple layer stack βœ… βœ…
One input βœ… βœ…
One output βœ… βœ…
Multiple inputs ❌ βœ…
Multiple outputs ❌ βœ…
Branching ❌ βœ…
Skip connections ❌ βœ…
Layer sharing ❌ βœ…
Complex graph ❌ βœ…
Easy to read βœ… βœ…
Complex architectures ❌ βœ…

πŸ— Branching ArchitectureΒΆ

Suppose an input should pass through two different paths.

flowchart TD

    INPUT["Input"]

    INPUT --> PATH1["Path 1<br>Dense 128"]

    INPUT --> PATH2["Path 2<br>Dense 64"]

    PATH1 --> MERGE["Concatenate"]

    PATH2 --> MERGE

    MERGE --> OUTPUT["Output"]

This cannot naturally be represented as a simple Sequential stack.

The Functional API handles it directly.


πŸ§ͺ Branching ExampleΒΆ

inputs = tf.keras.Input(
    shape=(100,)
)

branch_1 = tf.keras.layers.Dense(
    128,
    activation="relu"
)(inputs)

branch_2 = tf.keras.layers.Dense(
    64,
    activation="relu"
)(inputs)

merged = tf.keras.layers.Concatenate()(
    [
        branch_1,
        branch_2
    ]
)

outputs = tf.keras.layers.Dense(
    10,
    activation="softmax"
)(merged)

model = tf.keras.Model(
    inputs=inputs,
    outputs=outputs
)

🧠 Concatenation¢

Concatenation joins tensors along a specified dimension.

Conceptually:

Tensor A
[128 features]

        +

Tensor B
[64 features]

        ↓

Merged Tensor
[192 features]

The resulting tensor can then be passed to another layer.


πŸ”€ Add vs ConcatenateΒΆ

Two common merging operations are:

Add
Concatenate

AddΒΆ

Requires compatible shapes.

A ──┐
    β”œβ”€β”€ Add
B β”€β”€β”˜

ConcatenateΒΆ

Combines feature dimensions.

A ──┐
    β”œβ”€β”€ Concatenate
B β”€β”€β”˜

πŸ§ͺ Add ExampleΒΆ

x1 = tf.keras.layers.Dense(
    128
)(inputs)

x2 = tf.keras.layers.Dense(
    128
)(inputs)

merged = tf.keras.layers.Add()(
    [
        x1,
        x2
    ]
)

The two tensors must have compatible shapes.


🧠 Skip Connections¢

Skip connections allow information to bypass one or more layers.

flowchart TD

    INPUT["Input"]

    L1["Layer 1"]

    L2["Layer 2"]

    ADD["Add"]

    INPUT --> L1
    L1 --> L2
    L2 --> ADD

    INPUT --> ADD

    ADD --> OUTPUT["Output"]

Mathematically:

[ y=F(x)+x ]

This is a core idea behind residual networks.


πŸ§ͺ Skip Connection ExampleΒΆ

inputs = tf.keras.Input(
    shape=(128,)
)

x = tf.keras.layers.Dense(
    128,
    activation="relu"
)(inputs)

x = tf.keras.layers.Dense(
    128
)(x)

x = tf.keras.layers.Add()(
    [
        x,
        inputs
    ]
)

outputs = tf.keras.layers.Activation(
    "relu"
)(x)

model = tf.keras.Model(
    inputs=inputs,
    outputs=outputs
)

The Functional API makes this architecture straightforward.


🧠 Layer Reuse¢

Functional API allows the same layer to be reused.

For example:

shared_layer = tf.keras.layers.Dense(
    64,
    activation="relu"
)

The same layer can process multiple inputs.

x1 = shared_layer(input_1)
x2 = shared_layer(input_2)

This means the layer's weights are shared.


πŸ”„ Shared Layer ArchitectureΒΆ

flowchart TD

    INPUT1["Input 1"]
    INPUT2["Input 2"]

    SHARED["Shared Dense Layer"]

    INPUT1 --> SHARED
    INPUT2 --> SHARED

    SHARED --> OUT1["Output 1"]
    SHARED --> OUT2["Output 2"]

This pattern is useful for:

  • Siamese networks
  • Metric learning
  • Shared feature extraction
  • Multi-input architectures

πŸ§ͺ Multiple InputsΒΆ

Suppose a model receives:

Customer Profile
Transaction History

as separate inputs.

flowchart TD

    PROFILE["Customer Profile"]

    HISTORY["Transaction History"]

    PROFILE --> P["Profile Encoder"]

    HISTORY --> H["History Encoder"]

    P --> MERGE["Concatenate"]

    H --> MERGE

    MERGE --> OUTPUT["Risk Prediction"]

Functional API implementation:

profile_input = tf.keras.Input(
    shape=(20,),
    name="profile"
)

history_input = tf.keras.Input(
    shape=(50,),
    name="history"
)

profile_features = tf.keras.layers.Dense(
    64,
    activation="relu"
)(profile_input)

history_features = tf.keras.layers.Dense(
    128,
    activation="relu"
)(history_input)

merged = tf.keras.layers.Concatenate()(
    [
        profile_features,
        history_features
    ]
)

outputs = tf.keras.layers.Dense(
    1,
    activation="sigmoid"
)(merged)

model = tf.keras.Model(
    inputs=[
        profile_input,
        history_input
    ],
    outputs=outputs
)

🧠 Multiple Outputs¢

A model can also produce multiple outputs.

For example:

Input
  β”‚
  β”œβ”€β”€ Classification Head
  β”‚
  └── Regression Head
flowchart TD

    INPUT["Shared Input"]

    BACKBONE["Shared Feature Extractor"]

    INPUT --> BACKBONE

    BACKBONE --> CLASS["Classification Head"]

    BACKBONE --> REG["Regression Head"]

    CLASS --> CLASSOUT["Class Output"]

    REG --> REGOUT["Regression Output"]

πŸ§ͺ Multi-Output ModelΒΆ

inputs = tf.keras.Input(
    shape=(100,)
)

features = tf.keras.layers.Dense(
    128,
    activation="relu"
)(inputs)

class_output = tf.keras.layers.Dense(
    10,
    activation="softmax",
    name="classification"
)(features)

reg_output = tf.keras.layers.Dense(
    1,
    name="regression"
)(features)

model = tf.keras.Model(
    inputs=inputs,
    outputs=[
        class_output,
        reg_output
    ]
)

🧠 Multi-Task Learning¢

A multi-output model can support multi-task learning.

For example:

Shared Representation
        β”‚
   β”Œβ”€β”€β”€β”€β”΄β”€β”€β”€β”€β”
   ↓         ↓
Task A      Task B

One model can learn:

Classification
+
Regression

or:

Object Detection
+
Object Classification

or:

Sentiment
+
Topic Classification

βš™οΈ Compiling Multi-Output ModelsΒΆ

Different outputs can have different losses.

model.compile(

    optimizer="adam",

    loss={
        "classification":
            "sparse_categorical_crossentropy",

        "regression":
            "mse"
    },

    metrics={
        "classification":
            ["accuracy"],

        "regression":
            ["mae"]
    }
)

🧠 Loss Weighting¢

Multiple tasks may contribute differently to the total loss.

Conceptually:

[ L_{total} = \lambda_1L_1 + \lambda_2L_2 ]

Example:

model.compile(

    optimizer="adam",

    loss={
        "classification": "sparse_categorical_crossentropy",
        "regression": "mse"
    },

    loss_weights={
        "classification": 1.0,
        "regression": 0.5
    }
)

🧠 Model Composition¢

Keras models can themselves behave like layers.

For example:

encoder = tf.keras.Model(
    inputs=encoder_input,
    outputs=encoded
)

decoder = tf.keras.Model(
    inputs=decoder_input,
    outputs=decoded
)

They can then be composed:

autoencoder_output = decoder(
    encoder(inputs)
)

This is extremely useful for:

  • Autoencoders
  • Encoder-decoder architectures
  • Transfer Learning
  • Reusable model components

πŸ— Model-as-a-Layer ConceptΒΆ

flowchart LR

    INPUT["Input"]

    ENCODER["Encoder Model"]

    DECODER["Decoder Model"]

    OUTPUT["Output"]

    INPUT --> ENCODER
    ENCODER --> DECODER
    DECODER --> OUTPUT

🧠 Reusable Model Components¢

Instead of creating one huge model, break it into components:

Embedding
     ↓
Encoder
     ↓
Feature Extractor
     ↓
Task Head

This improves:

  • Reusability
  • Testing
  • Maintenance
  • Experimentation

πŸ§ͺ Building a Reusable BlockΒΆ

def dense_block(
    units,
    dropout_rate=0.2
):

    return tf.keras.Sequential([

        tf.keras.layers.Dense(
            units,
            activation="relu"
        ),

        tf.keras.layers.BatchNormalization(),

        tf.keras.layers.Dropout(
            dropout_rate
        )
    ])

Use:

block = dense_block(
    128
)

x = block(inputs)

🧠 Functional API and CNNs¢

The Functional API becomes particularly valuable for Computer Vision.

For example:

Input Image
     ↓
Convolution
     ↓
Pooling
     ↓
Branch
 β”Œβ”€β”€β”€β”΄β”€β”€β”€β”
 ↓       ↓
CNN     CNN
 β””β”€β”€β”€β”¬β”€β”€β”€β”˜
     ↓
 Merge
     ↓
Classifier

This allows architectures beyond simple sequential CNN stacks.


🧠 Functional API and ResNet¢

Residual networks rely heavily on skip connections.

Conceptually:

flowchart TD

    INPUT["Input"]

    CONV1["Convolution"]

    CONV2["Convolution"]

    ADD["Add"]

    RELU["ReLU"]

    INPUT --> CONV1
    CONV1 --> CONV2
    CONV2 --> ADD

    INPUT --> ADD

    ADD --> RELU

The Functional API is naturally suited to this type of architecture.

ResNet is covered in detail in:

22. ResNet, Residual Connections and TorchVision


🧠 Functional API and Siamese Networks¢

Siamese networks often process two inputs through shared weights.

flowchart TD

    INPUT1["Input A"]

    INPUT2["Input B"]

    SHARED["Shared Encoder"]

    INPUT1 --> SHARED
    INPUT2 --> SHARED

    SHARED --> EMB1["Embedding A"]
    SHARED --> EMB2["Embedding B"]

    EMB1 --> DIST["Distance / Similarity"]

    EMB2 --> DIST

The Functional API is a natural fit because the same encoder can be reused.


🧠 Functional API and Attention¢

Attention architectures commonly contain:

Query
Key
Value

and multiple computation paths.

Functional graphs can represent such architectures more naturally than simple sequential stacks.

This becomes increasingly important when building:

  • Attention models
  • Transformers
  • Encoder-decoder systems
  • Multi-head architectures

🧠 Functional API Model Visualization¢

Keras can visualize the architecture.

tf.keras.utils.plot_model(
    model,
    show_shapes=True,
    show_layer_names=True
)

For complex architectures, this is extremely useful for verifying:

Input Shapes
Output Shapes
Connections
Branches
Layer Names

🧠 Model Summary¢

model.summary()

For a Functional model, the summary can reveal:

Layer
Output Shape
Parameters
Connections

This helps identify:

  • Unexpected parameter growth
  • Shape mismatches
  • Incorrect architecture
  • Large layers

πŸ” Inspecting the GraphΒΆ

Functional models expose:

model.inputs
model.outputs
model.layers

Example:

for layer in model.layers:

    print(
        layer.name,
        layer.output.shape
    )

🧠 Naming Layers¢

Meaningful layer names improve observability.

Instead of:

tf.keras.layers.Dense(128)

use:

tf.keras.layers.Dense(
    128,
    activation="relu",
    name="feature_projection"
)

This becomes especially useful in:

  • Debugging
  • Model visualization
  • Monitoring
  • Transfer Learning
  • Model inspection

🏒 Production Architecture¢

A production Keras model should ideally have clear boundaries between:

Input Contract
      ↓
Preprocessing
      ↓
Feature Extraction
      ↓
Task-Specific Head
      ↓
Output Contract

Example:

flowchart LR

    INPUT["Input Contract"]

    PRE["Preprocessing"]

    FEATURES["Feature Extractor"]

    HEAD["Prediction Head"]

    OUTPUT["Output Contract"]

    INPUT --> PRE
    PRE --> FEATURES
    FEATURES --> HEAD
    HEAD --> OUTPUT

🧠 Functional API as Architecture-as-Code¢

The Functional API allows the model topology to be expressed directly in Python.

For example:

inputs = tf.keras.Input(
    shape=(128,)
)

x = layer_a(inputs)

x1 = layer_b(x)

x2 = layer_c(x)

merged = tf.keras.layers.Add()(
    [x1, x2]
)

outputs = layer_d(
    merged
)

model = tf.keras.Model(
    inputs,
    outputs
)

The code itself reflects the architecture:

Input
  ↓
Layer A
  ↓
 β”Œβ”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”
 ↓             ↓
Layer B       Layer C
 β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
        ↓
       Add
        ↓
     Layer D
        ↓
      Output

🧠 Sequential β†’ Functional MigrationΒΆ

A Sequential model:

model = tf.keras.Sequential([

    tf.keras.Input(
        shape=(784,)
    ),

    tf.keras.layers.Dense(
        128,
        activation="relu"
    ),

    tf.keras.layers.Dense(
        10,
        activation="softmax"
    )
])

can be expressed using Functional API:

inputs = tf.keras.Input(
    shape=(784,)
)

x = tf.keras.layers.Dense(
    128,
    activation="relu"
)(inputs)

outputs = tf.keras.layers.Dense(
    10,
    activation="softmax"
)(x)

model = tf.keras.Model(
    inputs,
    outputs
)

The underlying neural network is equivalent.

The difference is the amount of architectural control.


🧠 When Should You Use Sequential?¢

Use Sequential when:

One Input
+
One Output
+
Linear Stack
+
No Branching
+
No Layer Sharing
+
No Skip Connections

Example:

Input
 ↓
Dense
 ↓
Dropout
 ↓
Dense
 ↓
Output

🧠 When Should You Use Functional API?¢

Use Functional API when you need:

Multiple Inputs
Multiple Outputs
Branches
Merges
Skip Connections
Shared Layers
Complex Graphs
Reusable Components

Examples:

ResNet
Siamese Networks
Multi-Task Models
Encoder-Decoder Models
Complex CNNs
Attention Networks
Transformers

🧠 When Should You Use Model Subclassing?¢

Model subclassing becomes useful when:

  • The architecture is highly dynamic
  • Forward behavior requires custom control flow
  • You need custom training behavior
  • The computation graph cannot be conveniently expressed using standard Functional patterns

Example:

class MyModel(
    tf.keras.Model
):

    def __init__(self):

        super().__init__()

        self.dense = tf.keras.layers.Dense(
            128,
            activation="relu"
        )

        self.output_layer = tf.keras.layers.Dense(
            10
        )

    def call(
        self,
        inputs
    ):

        x = self.dense(
            inputs
        )

        return self.output_layer(
            x
        )

Model subclassing is covered further in:

15. Custom Layers, Models and Training Loops


🧭 Model API Decision Guide¢

flowchart TD

    START["Choose Keras Model API"]

    START --> SIMPLE{"Simple Linear Stack?"}

    SIMPLE -->|Yes| SEQ["Sequential API"]

    SIMPLE -->|No| COMPLEX{"Complex Graph?"}

    COMPLEX -->|Yes| FUNC["Functional API"]

    COMPLEX -->|No| CUSTOM{"Custom Dynamic Behavior?"}

    CUSTOM -->|Yes| SUB["Model Subclassing"]

    CUSTOM -->|No| FUNC

πŸ§ͺ Practical Example β€” Multi-Input Enterprise ModelΒΆ

Imagine a financial risk model receiving:

Customer Profile
+
Transaction Features
+
Behavioral Features

Architecture:

flowchart TD

    PROFILE["Customer Profile"]

    TRANS["Transaction Features"]

    BEHAVIOR["Behavioral Features"]

    PROFILE --> PENC["Profile Encoder"]

    TRANS --> TENC["Transaction Encoder"]

    BEHAVIOR --> BENC["Behavior Encoder"]

    PENC --> MERGE["Feature Fusion"]

    TENC --> MERGE

    BENC --> MERGE

    MERGE --> DENSE["Dense Layers"]

    DENSE --> RISK["Risk Score"]

This is an excellent use case for the Functional API.


πŸ§ͺ ImplementationΒΆ

import tensorflow as tf


profile = tf.keras.Input(
    shape=(20,),
    name="profile"
)

transactions = tf.keras.Input(
    shape=(50,),
    name="transactions"
)

behavior = tf.keras.Input(
    shape=(30,),
    name="behavior"
)


profile_features = tf.keras.layers.Dense(
    64,
    activation="relu",
    name="profile_encoder"
)(profile)


transaction_features = tf.keras.layers.Dense(
    128,
    activation="relu",
    name="transaction_encoder"
)(transactions)


behavior_features = tf.keras.layers.Dense(
    64,
    activation="relu",
    name="behavior_encoder"
)(behavior)


features = tf.keras.layers.Concatenate(
    name="feature_fusion"
)([
    profile_features,
    transaction_features,
    behavior_features
])


x = tf.keras.layers.Dense(
    128,
    activation="relu",
    name="risk_features"
)(features)


x = tf.keras.layers.Dropout(
    0.2,
    name="regularization"
)(x)


output = tf.keras.layers.Dense(
    1,
    activation="sigmoid",
    name="risk_score"
)(x)


model = tf.keras.Model(
    inputs=[
        profile,
        transactions,
        behavior
    ],
    outputs=output,
    name="enterprise_risk_model"
)

🧠 Why This Architecture Is Production-Friendly¢

Each input has a dedicated representation:

Profile
   ↓
Profile Encoder

Transactions
   ↓
Transaction Encoder

Behavior
   ↓
Behavior Encoder

These representations are then combined:

Feature Fusion
      ↓
Shared Representation
      ↓
Prediction Head

This structure makes it easier to:

  • Modify individual branches
  • Test components
  • Reuse encoders
  • Monitor inputs
  • Extend the architecture

⚠ Common Mistakes¢

Avoid these mistakes:

  • Using Sequential for a graph-based architecture
  • Using Functional API when Sequential would be simpler without a reason
  • Forgetting that Functional API tensors are symbolic during model construction
  • Connecting tensors with incompatible shapes
  • Using Add when tensor shapes are incompatible
  • Confusing Add with Concatenate
  • Accidentally creating separate layers instead of sharing the same layer instance
  • Creating unnecessarily complicated model graphs
  • Not naming important inputs and outputs
  • Ignoring model visualization
  • Ignoring parameter counts
  • Building giant monolithic Functional models
  • Mixing preprocessing logic inconsistently between training and inference
  • Forgetting to validate multi-input data contracts
  • Using different preprocessing logic for different deployment paths
  • Adding branches without a clear modeling reason

🧠 Interview Questions¢

BeginnerΒΆ

1. What is the Sequential API?ΒΆ

The Sequential API is a Keras model-building approach for creating a simple linear stack of layers.

2. What is the Functional API?ΒΆ

The Functional API allows developers to construct arbitrary directed computation graphs by explicitly connecting inputs, layers, and outputs.

3. What is the main difference between Sequential and Functional API?ΒΆ

Sequential represents a simple linear stack, while Functional API supports complex graph structures such as branching, merging, multiple inputs, multiple outputs, shared layers, and skip connections.

4. Can Sequential models have multiple inputs?ΒΆ

Not naturally. Models with multiple inputs should generally use the Functional API.

5. Can Functional API build simple models?ΒΆ

Yes. A simple Sequential model can also be represented using the Functional API.


IntermediateΒΆ

6. What is a symbolic tensor?ΒΆ

A symbolic tensor represents a node or intermediate value in the Keras computation graph rather than an actual batch of numerical data during model construction.

7. What is a skip connection?ΒΆ

A skip connection allows information to bypass one or more layers and later be combined with the transformed representation.

8. Why is the Functional API useful for ResNet?ΒΆ

ResNet uses residual/skip connections, which require graph structures that are not naturally represented by a simple sequential stack.

9. What is layer sharing?ΒΆ

Layer sharing means using the same layer instance, and therefore the same learned weights, on multiple inputs or paths.

10. What is a multi-input model?ΒΆ

A model that receives more than one independent input tensor.

11. What is a multi-output model?ΒΆ

A model that produces multiple outputs, potentially representing different tasks.

12. What is multi-task learning?ΒΆ

Multi-task learning trains a shared representation to solve multiple related tasks, often using separate task-specific output heads.


AdvancedΒΆ

13. Why would you use the Functional API instead of subclassing?ΒΆ

When the architecture can be clearly expressed as a computation graph, the Functional API provides strong graph visibility, model inspection, visualization, and serialization while remaining flexible.

14. Why might model subclassing be necessary?ΒΆ

Highly dynamic computation, custom control flow, or specialized training behavior may be easier to implement using subclassing.

15. What is the difference between Add and Concatenate?ΒΆ

Add performs element-wise addition and requires compatible shapes. Concatenate joins tensors along a specified dimension and generally increases the feature dimension.

16. How does layer sharing work?ΒΆ

The same layer object is applied to multiple inputs:

shared_layer = Dense(64)

x1 = shared_layer(input1)
x2 = shared_layer(input2)

Both paths use the same learned weights.

17. How would you design a multi-input enterprise model?ΒΆ

Separate each input into an appropriate encoder, transform each representation, fuse the representations, and pass the fused representation into shared prediction layers.

18. Why is model visualization important?ΒΆ

It helps verify:

Connections
Shapes
Branches
Skip Paths
Parameter Growth

and can expose architectural mistakes before training becomes expensive.


πŸ§ͺ Practical ExercisesΒΆ

Exercise 1 β€” Sequential ClassifierΒΆ

Build:

Input
 ↓
Dense 128
 ↓
ReLU
 ↓
Dense 64
 ↓
ReLU
 ↓
Dense 10
 ↓
Softmax

using Sequential.


Exercise 2 β€” Convert Sequential to FunctionalΒΆ

Implement the same architecture using the Functional API.

Compare:

Code
Model Summary
Parameter Count
Predictions

Exercise 3 β€” Branching NetworkΒΆ

Create:

Input
 β”œβ”€β”€ Dense 128
 β”‚
 └── Dense 64
       ↓
   Concatenate
       ↓
   Dense
       ↓
   Output

Exercise 4 β€” Skip ConnectionΒΆ

Implement:

Input
   β”‚
   β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
   ↓               β”‚
Dense              β”‚
   ↓               β”‚
Dense              β”‚
   ↓               β”‚
   └──── Add β—„β”€β”€β”€β”€β”€β”˜
          ↓
        Output

Exercise 5 β€” Multi-Input ModelΒΆ

Build a model accepting:

Profile Features
Transaction Features
Behavior Features

Create separate encoders and combine them using Concatenate.


Exercise 6 β€” Multi-Output ModelΒΆ

Create one shared network with:

Classification Head
+
Regression Head

Train it using different losses for each output.


Exercise 7 β€” Shared EncoderΒΆ

Build a Siamese-style architecture:

Input A ──┐
          β”œβ”€β”€ Shared Encoder
Input B β”€β”€β”˜

Produce two embeddings and calculate a similarity score.


πŸ“Œ Key TakeawaysΒΆ

  • Keras provides multiple model-construction approaches.
  • Sequential is ideal for simple linear stacks.
  • Functional API is designed for arbitrary computation graphs.
  • Functional API supports multiple inputs and outputs.
  • Functional API supports branching and merging.
  • Functional API supports shared layers.
  • Skip connections are naturally implemented using Functional API.
  • Residual networks rely heavily on this capability.
  • Functional API uses symbolic tensors during model construction.
  • tf.keras.Model(inputs, outputs) creates a Functional model.
  • Add performs element-wise addition.
  • Concatenate joins tensors along a dimension.
  • Multi-output models can support multi-task learning.
  • Different outputs can use different losses and metrics.
  • Loss weights can balance multiple tasks.
  • Keras models can be composed and reused as layers.
  • Meaningful layer names improve model inspection and maintainability.
  • Model visualization is especially valuable for complex architectures.
  • Sequential should not be forced onto architectures that require graph connectivity.
  • Functional API is often the right abstraction for production-grade complex neural networks.
  • Model subclassing is useful when computation or behavior requires greater customization.
  • Good architecture design should balance flexibility, readability, maintainability, and operational requirements.

πŸ“š Further ReadingΒΆ

Continue with:

The next chapter moves from standard Keras model construction into custom layers, custom models, and custom training loops, giving you much deeper control over the training process.


➑️ Next Chapter¢

15. Custom Layers, Models and Training Loops


Enterprise AI Engineering Handbook
Building Production-Grade Enterprise AI Systems β€” One Chapter at a Time.