C.W.K.
Stream
Lesson 03 of 04 · published

HParams Dashboard and the Profiler

~12 min · hparams, profiler, performance

Level 0Level 0
0 XP0/78 lessons0/17 achievements
0/100 XP to next level100 XP to go0% complete

Find good hyperparameters and find your bottleneck

The HParams dashboard compares dozens of training runs with different hyperparameters. Define a search space, run experiments logging each run's hyperparameters and final metrics, then explore parallel coordinate plots, scatter plots, and sortable tables to find what works.

The TensorBoard Profiler identifies performance bottlenecks. Most TF performance issues are not slow compute — they're slow data loading. The Profiler's Input Pipeline Analyzer tells you exactly what fraction of time is spent waiting for data vs running model computation. "Input Bound: 80%" means your tf.data pipeline needs work, not your model.

Code

HParams sweep·python
from tensorboard.plugins.hparams import api as hp
import tensorflow as tf

# Define the search space
HP_NUM_UNITS = hp.HParam('num_units',  hp.Discrete([32, 64, 128]))
HP_DROPOUT   = hp.HParam('dropout',    hp.RealInterval(0.1, 0.5))
HP_OPTIMIZER = hp.HParam('optimizer',  hp.Discrete(['adam', 'sgd']))
METRIC_ACC   = 'accuracy'

with tf.summary.create_file_writer('logs/hparam_tuning').as_default():
    hp.hparams_config(
        hparams=[HP_NUM_UNITS, HP_DROPOUT, HP_OPTIMIZER],
        metrics=[hp.Metric(METRIC_ACC, display_name='Validation Accuracy')],
    )

def train_test_model(hparams):
    model = tf.keras.Sequential([
        tf.keras.layers.Dense(hparams[HP_NUM_UNITS], activation='relu',
                              input_shape=(784,)),
        tf.keras.layers.Dropout(hparams[HP_DROPOUT]),
        tf.keras.layers.Dense(10, activation='softmax'),
    ])
    model.compile(optimizer=hparams[HP_OPTIMIZER],
                  loss='sparse_categorical_crossentropy',
                  metrics=['accuracy'])
    model.fit(x_train, y_train, epochs=5, verbose=0)
    _, accuracy = model.evaluate(x_test, y_test, verbose=0)
    return accuracy

session_num = 0
for num_units in HP_NUM_UNITS.domain.values:
    for dropout in [0.1, 0.3, 0.5]:
        for optimizer in HP_OPTIMIZER.domain.values:
            hparams = {
                HP_NUM_UNITS: num_units, HP_DROPOUT: dropout,
                HP_OPTIMIZER: optimizer,
            }
            run_dir = f"logs/hparam_tuning/run-{session_num:03d}"
            with tf.summary.create_file_writer(run_dir).as_default():
                hp.hparams(hparams)
                accuracy = train_test_model(hparams)
                tf.summary.scalar(METRIC_ACC, accuracy, step=1)
            session_num += 1
Profiler — 병목 찾기·python
# pip install tensorboard_plugin_profile
import tensorflow as tf

# Method 1: profile a batch range via callback
tb_callback = tf.keras.callbacks.TensorBoard(
    log_dir='logs/profile',
    profile_batch='5, 15',     # profile batches 5 through 15
)
model.fit(train_ds, epochs=3, callbacks=[tb_callback])

# Method 2: programmatic
tf.profiler.experimental.start('logs/profile')
model.fit(train_ds, epochs=1)
tf.profiler.experimental.stop()

# Method 3: context manager
with tf.profiler.experimental.Profile('logs/profile'):
    model.fit(train_ds, epochs=1)

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign inPlease sign in to comment.

No comments yet — be the first.