Your First Pipeline
In etl4s, everything is either:
- A
Node[-In, +Out] - A
Nodewrapped in a Reader monad (Reader[Cfg, Node[In, Out]])Cfgis the configuration type needed to run the node
Nodes are simple lambdas from In => Out
You can run them like functions or be more deliberate with .unsafeRun(In) ... just
a matter of taste.
Create new nodes by chaining existing ones together:
Here is a more substantive example:
val extract5 = Node(5)
val timesTwo = Node[Int, Int](_ * 2)
val consoleLoad = Node[Int, Unit](x => println(s"Result: $x"))
val p =
extract5 ~> timesTwo ~> consoleLoad
p.unsafeRun()
This will give:
To improve readability and express intent, etl4s defines three aliases: Extract, Transform and Load. All behave the same under the hood.
type Extract[In, Out] = Node[In, Out]
type Transform[In, Out] = Node[In, Out]
type Load[In, Out] = Node[In, Out]
You can use other operators like & to fan out and stitch your graphs
val double = Node[Int, Int](_ * 2)
val triple = Node[Int, Int](_ * 3)
val combine = Node[(Int, Int), Int] { case (a, b) => a + b }
val p =
extract5 ~> (double & triple) ~> combine
p.unsafeRun() // 25
One of the key benefits of etl4s is that you can separate configuration (the "knobs to turn") from your actual flow of data.
val YEAR = 2025
val loadData = Node[Any, String] { _ =>
println(s"Loading $YEAR data")
"TEST DATA"
}
loadData.unsafeRun()
Prints Loading 2025 data and returns "TEST DATA".
That works for one value, but it doesn't compose: every node that reaches for year is
an invisible, untyped dependency.
etl4s makes the knob an explicit input instead, declare it
with .requires, then .provide it once at the edge:
case class Config(year: Int)
val loadData = Node[Unit, String].requires[Config] {
config => _ => s"Loading ${config.year} data"
}
loadData.provide(Config(2025)).unsafeRun(()) // "Loading 2025 data"
.requires turns the node into a Reader[Config, Node[...]], and config-aware and plain nodes
compose together with the same ~>. See Configuration for details.
Pipelines are values
Building a pipeline runs nothing. Every combinator (~>, &, >>, ...) just grows an
immutable AST - a free profunctor over your plain functions.
Because a pipeline is just this tree, you can interpret it however you like. That is what
makes etl4s effect polymorphic: .compile[F] folds the same tree into In => F[Out] for
any effect F (Try, Future, cats-effect IO, ZIO, Kyo ...). See
Effect polymorphism.
Inspect the structure
Since a pipeline is a value, you can look at it before running it.
Every Node carries its own shape, its in/out types and the enclosing val name,
both captured at compile time by a small macro, so you can dump its stages or render it as a diagram:
Take the fan-out / fan-in pipeline from earlier and add a load step that writes the result:
.toDot renders a Graphviz graph, and .toMermaid a Mermaid one.
You get:
Both take options: showTypes = false drops the type labels on the edges, and direction
changes the layout (Direction.LR, TB, RL, BT):
.stages gives you the steps in execution order, each with its val-name and in/out types:
etl4s