January, 2024

article thumbnail

The Future of Data Engineering as a Data Engineer

Monte Carlo

In the world of data engineering, Maxime Beauchemin is someone who needs no introduction. One of the first data engineers at Facebook and Airbnb, he wrote and open sourced the wildly popular orchestrator, Apache Airflow , followed shortly thereafter by Apache Superset , a data exploration tool that’s taking the data viz landscape by storm. Currently, Maxime is CEO and co-founder of Preset , a fast-growing startup that’s paving the way forward for AI-enabled data visualization for modern companie

article thumbnail

The Only Free Course You Need To Become a Professional Data Engineer

KDnuggets

Data Engineering ZoomCamp offers free access to reading materials, video tutorials, assignments, homeworks, projects, and workshops.

article thumbnail

Data Science vs Software Engineering - Significant Differences

Knowledge Hut

With an array of career options, all that matters is choosing the right career path. The right career path for one depends on their skill set, interest, job availability in that field, and, most importantly, your passion for the same. Speaking of job vacancies, the two careers have high demands till date and in upcoming years are Data Scientist and a Software Engineer.

article thumbnail

Build A Data Lake For Your Security Logs With Scanner

Data Engineering Podcast

Summary Monitoring and auditing IT systems for security events requires the ability to quickly analyze massive volumes of unstructured log data. The majority of products that are available either require too much effort to structure the logs, or aren't fast enough for interactive use cases. Cliff Crosland co-founded Scanner to provide fast querying of high scale log data for security auditing.

Data Lake 147
article thumbnail

Apache Airflow® Best Practices for ETL and ELT Pipelines

Whether you’re creating complex dashboards or fine-tuning large language models, your data must be extracted, transformed, and loaded. ETL and ELT pipelines form the foundation of any data product, and Airflow is the open-source data orchestrator specifically designed for moving and transforming data in ETL and ELT pipelines. This eBook covers: An overview of ETL vs.

article thumbnail

LLM Training and Inference with Intel(R) Gaudi(R) 2 AI Accelerators

databricks

At Databricks, we want to help our customers build and deploy generative AI applications on their own data without sacrificing data privacy or.

Building 145
article thumbnail

Totally Eclipsed

ArcGIS

Exploring the value of critique as part of the process of creating a new map of the Total Eclipse that will cross the United States on April 8th

Process 143

More Trending

article thumbnail

AI Prompt Engineers are Making $300k/y

KDnuggets

Prompt engineering and generative AI are becoming hotter by the day. Be part of the heat!

article thumbnail

Robinhood Adds New Spot Bitcoin ETFs

Robinhood

The new class of spot Bitcoin ETFs that were approved by the SEC yesterday are now available on Robinhood Earlier today, Robinhood started offering the new class of spot Bitcoin ETFs that were approved by the SEC on January 10. These 11 ETFs became tradable to all customers in the United States this morning in both retirement and brokerage accounts though Robinhood Financial.

Insurance 132
article thumbnail

Modern Customer Data Platform Principles

Data Engineering Podcast

Summary Databases and analytics architectures have gone through several generational shifts. A substantial amount of the data that is being managed in these systems is related to customers and their interactions with an organization. In this episode Tasso Argyros, CEO of ActionIQ, gives a summary of the major epochs in database technologies and how he is applying the capabilities of cloud data warehouses to the challenge of building more comprehensive experiences for end-users through a modern c

Data Lake 147
article thumbnail

Welcome to the Data Intelligence Platform: Databricks + Einblick

databricks

At Databricks, we believe that AI will change the way that enterprises interact with their data. That’s why today, we're excited to welcome t.

Data 139
article thumbnail

Apache Airflow®: The Ultimate Guide to DAG Writing

Speaker: Tamara Fingerlin, Developer Advocate

In this new webinar, Tamara Fingerlin, Developer Advocate, will walk you through many Airflow best practices and advanced features that can help you make your pipelines more manageable, adaptive, and robust. She'll focus on how to write best-in-class Airflow DAGs using the latest Airflow features like dynamic task mapping and data-driven scheduling!

article thumbnail

Geoprocessing enhancements in ArcGIS Pro 3.2 for ArcMap users

ArcGIS

Equivalency enhancements to geoprocessing in ArcGIS Pro 3.2 to remove more barriers for those transitioning from ArcMap.

139
139
article thumbnail

Apache Flink and cluster components deep dive

Waitingforcode

Previously you could read about transformation of a user job definition into an executable stream graph. Since this explanation was relatively high-level, I decided to deep dive into the final step executing the code.

Coding 130
article thumbnail

Top 16 Technical Data Sources for Advanced Data Science Projects

KDnuggets

Here are data repositories that will up your data science game and improve your data projects.

article thumbnail

The State of Data Engineering at Data Day Texas 2024

Jesse Anderson

The premier of my latest talk covering The State of Data Engineering. I go through the history of the industry to see where we’re heading. This starts with data warehousing and goes into data science. I finish off by showing how data engineering can avoid the same fate as data warehousing and data science. Sorry, we didn’t have a microphone for the questions and I forgot to repeat some of the questions.

article thumbnail

Optimizing The Modern Developer Experience with Coder

Many software teams have migrated their testing and production workloads to the cloud, yet development environments often remain tied to outdated local setups, limiting efficiency and growth. This is where Coder comes in. In our 101 Coder webinar, you’ll explore how cloud-based development environments can unlock new levels of productivity. Discover how to transition from local setups to a secure, cloud-powered ecosystem with ease.

article thumbnail

Pushing The Limits Of Scalability And User Experience For Data Processing WIth Jignesh Patel

Data Engineering Podcast

Summary Data processing technologies have dramatically improved in their sophistication and raw throughput. Unfortunately, the volumes of data that are being generated continue to double, requiring further advancements in the platform capabilities to keep up. As the sophistication increases, so does the complexity, leading to challenges for user experience.

article thumbnail

Databricks Announces the Industry’s First Generative AI Engineer Learning Pathway and Certification

databricks

Today, we are announcing the industry's first Generative AI Engineer learning pathway and certification to help ensure that data and AI practitioners have.

article thumbnail

Introducing Neighborhood Explorer in ArcGIS Pro

ArcGIS

ArcGIS Pro now includes Neighborhood Explorer: an experience that will help you understand and refine spatial relationships in your analysis.

Education 138
article thumbnail

Data News — Week 24.04

Christophe Blefari

Hey ( credits ) Hey, new week new email. This is already end of January but I took time to travel and see people I did not see for a long time so I'm super happy how this new year is starting. Next week, I'll be wrapping up my DataOps lecture by incorporating how to deploy machine learning models. This is a fun part where students learn how to serve a simple classifier in production.

Algorithm 130
article thumbnail

15 Modern Use Cases for Enterprise Business Intelligence

Large enterprises face unique challenges in optimizing their Business Intelligence (BI) output due to the sheer scale and complexity of their operations. Unlike smaller organizations, where basic BI features and simple dashboards might suffice, enterprises must manage vast amounts of data from diverse sources. What are the top modern BI use cases for enterprise businesses to help you get a leg up on the competition?

article thumbnail

7 Steps to Landing Your First Data Science Job

KDnuggets

Want to make a successful career switch to data science? From learning data science concepts to cracking interviews, read this guide to move one step closer to your first data science job.

article thumbnail

Static enrichment dataset with Delta Lake

Waitingforcode

Data enrichment is one of common data engineering tasks. It's relatively easy to implement with static datasets because of the data availability. However, this apparently easy task can become a nightmare if used with inappropriate technologies.

Datasets 130
article thumbnail

Cutting Your Data Stack Costs: How To Approach It And Common Issues

Seattle Data Guy

I once had an engineer tell me that they essentially didn’t want to consider cost as they were building a solution. I was baffled. Don’t get me wrong, yes, when you’re building, you iterate and aim to improve your solutions cost. But from my perspective, I don’t think completely ignoring costs from day one is… Read more The post Cutting Your Data Stack Costs: How To Approach It And Common Issues appeared first on Seattle Data Guy.

IT 130
article thumbnail

Databricks SQL Year in Review (Part I): AI-optimized Performance and Serverless Compute

databricks

This is part 1 of a blog series where we look back at the major areas of progress for Databricks SQL in 2023.

SQL 138
article thumbnail

Prepare Now: 2025s Must-Know Trends For Product And Data Leaders

Speaker: Jay Allardyce, Deepak Vittal, Terrence Sheflin, and Mahyar Ghasemali

As we look ahead to 2025, business intelligence and data analytics are set to play pivotal roles in shaping success. Organizations are already starting to face a host of transformative trends as the year comes to a close, including the integration of AI in data analytics, an increased emphasis on real-time data insights, and the growing importance of user experience in BI solutions.

article thumbnail

Cartographic conventions

ArcGIS

What are cartographic conventions and do you need to follow them?

Education 129
article thumbnail

How to learn data engineering

Christophe Blefari

Learn data engineering, all the references ( credits ) This is a special edition of the Data News. But right now I'm in holidays finishing a hiking week in Corsica 🥾 So I wrote this special edition about: how to learn data engineering in 2024. The aim of this post is to create a repository of important links and concepts we should care about when we do data engineering.

article thumbnail

4 Steps to Become a Generative AI Developer

KDnuggets

In this post, we will cover what a generative AI developer does, what tools you need to master, and how to get started.

153
153
article thumbnail

Table file formats - streaming reader: Delta Lake

Waitingforcode

Even though I'm into streaming these days, I haven't really covered streaming in Delta Lake yet. I only slightly blogged about Change Data Feed but completely missed the fundamentals. Hopefully, this and next blog posts will change this!

Data 130
article thumbnail

How to Drive Cost Savings, Efficiency Gains, and Sustainability Wins with MES

Speaker: Nikhil Joshi, Founder & President of Snic Solutions

Is your manufacturing operation reaching its efficiency potential? A Manufacturing Execution System (MES) could be the game-changer, helping you reduce waste, cut costs, and lower your carbon footprint. Join Nikhil Joshi, Founder & President of Snic Solutions, in this value-packed webinar as he breaks down how MES can drive operational excellence and sustainability.

article thumbnail

7 Great Embedded Analytics Solutions – Which Embedded Analytics Solutions Should You Use?

Seattle Data Guy

Big data is big business these days. Organizations that hope to get ahead in crowded markets must utilize data from a variety of often highly disparate sources to understand how they’re performing and what customers are saying about them. However, data without the right analysis and reporting tools is just a waste of digital storage… Read more The post 7 Great Embedded Analytics Solutions – Which Embedded Analytics Solutions Should You Use?

Big Data 130
article thumbnail

Serving Quantized LLMs on NVIDIA H100 Tensor Core GPUs

databricks

Quantization is a technique for making machine learning models smaller and faster. We quantize Llama2-70B-Chat, producing an equivalent-quality model that generates 2.2x more.

article thumbnail

Polars vs Spark

Confessions of a Data Guy

The post Polars vs Spark appeared first on Confessions of a Data Guy.

Data 130
article thumbnail

Data News — Week 24.03

Christophe Blefari

Walking in the street be like recently ( credits ) Hey I hope this new edition finds you well. We are deep in the winter, it's time for comfy Data News to read near the fire 🔥 This week, on Monday, I started my annual university lecture. It's been 9 years since I started teaching and this year something was different. The students were incredibly calm, obviously my course is a bit difficult at the beginning because it touches on concepts that they are not used to—cloud,

Data 130
article thumbnail

Improving the Accuracy of Generative AI Systems: A Structured Approach

Speaker: Anindo Banerjea, CTO at Civio & Tony Karrer, CTO at Aggregage

When developing a Gen AI application, one of the most significant challenges is improving accuracy. This can be especially difficult when working with a large data corpus, and as the complexity of the task increases. The number of use cases/corner cases that the system is expected to handle essentially explodes. 💥 Anindo Banerjea is here to showcase his significant experience building AI/ML SaaS applications as he walks us through the current problems his company, Civio, is solving.