This is quite common practice for at least Spark, but it may depend more on the database selection. Like Cassandra is very fast when doing lookup by full primary key (if you don't have a lot of data, enabling row cache may also help), I think that Couchbase may also be a good choice (I didn't work with it for a long time). Code may look something like this (full code is in this Zeppelin Notebook. This code also requires Spark Cassandra Connector 2.5.0, that has support for "direct join with Cassandra" - see the blog post on that release):
val streamingInputDF = spark.readStream
.format("kafka")
.option("kafka.bootstrap.servers", "10.101.36.9:9092")
.option("subscribe", "tweets2")
.load()
val tweetDF = streamingInputDF.selectExpr("CAST(value AS STRING)")
.select(from_json($"value", schema).as("tweet"))
.select($"tweet.payload.created_at".as("created_at").cast(TimestampType),
$"tweet.payload.lang".as("language"))
val streamingCountsDF = tweetDF
.where(col("language").isNotNull)
.groupBy($"language", window($"created_at", "1 minutes"))
.count()
.select($"language", $"window.start".as("ts"), $"count")
import org.apache.spark.sql.cassandra._
val lang_details = spark.read.cassandraFormat("languages", "zep").load()
val joined = streamingCountsDF.join(lang_details,
lang_details("id") === streamingCountsDF("language"), "left_outer")
.select($"language", $"native_name".as("lang_name"), $"ts", $"count")
...
Also, you need to take into account other requirements - how often you'll get updates, how fast these updates need to be propagated to streaming job, etc. For Spark, you may, for example, have a separate dataframe for data that is stored in the database, and you cache this data for faster joining, but refresh dataframe every N minutes, to get latest updates from DB. (you can find source code for that in the Stream Processing with Apache Spark book)