Distributed Thoughts
Keeping you hard wired to advances in all things distributed: from engineering systems at scale — to reliably moving data for our AI overlords.
- Indexed issues, last 90 days
- 2
- Latest publication
- Aug 24, 2026
- Audience
- Checking…
- Earliest in this view
- Aug 12, 2026
Latest issues
Advancing the Open Lakehouse with Apache Spark, the Delta Kernel, and the new UC Delta APIs (opens the original)
Read excerpt
<img alt="" class="sizing-normal" height="794" src="https://substackcdn.com/image/fetch/$s_!2ArS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b94b8df-b30c-430d-994f-e520c5988226_2816x1536.j
Working with Apache Spark DataFrames, JSON, and the good ol’ StructType (opens the original)
Read excerpt
This post is from 2018, was published on my Medium.So I have been lucky enough to work with Apache Spark for the last two years and in the countless projects I work on I find that there are usually many ways of doing the same thing, and sometimes things seem so easy but unfortunately the documentation will not cover all of your potential use cases. I am going
Publishing over time
Last 90 days. Choose a month to open its work.
Recurring subjects
Named in the text we hold. One piece can cover several.
Audience
No verified audience measurement yet.
About this data
Counts cover the work we have indexed. Tone needs enough text and a confident classification. Excerpts and episode notes are not full articles or transcripts.
Identity or attribution wrong? Suggest a correction.