Top Data Engineering Digest Cloud Computing Deep Learning Content for Week of Nov 16

Sat.Nov 16, 2019 - Fri.Nov 22, 2019

Advice for New and Junior Data Scientists

KDnuggets

NOVEMBER 22, 2019

If you are a new Data Scientists early in your professional journey, and you’re a bit confused and lost, then follow this advice to figure out how to best contribute to your company.

Data

Introducing ksqlDB

Confluent

NOVEMBER 20, 2019

Today marks a new release of KSQL, one so significant that we’re giving it a new name: ksqlDB. Like KSQL, ksqlDB remains freely available and community licensed, and you can […].

IT Process Management

Join 37,000+

Insiders

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Webinars

Agent Tooling: Connecting AI to Your Tools, Systems & Data

How to Modernize Manufacturing Without Losing Control

Mastering Apache Airflow® 3.0: What’s New (and What’s Next) for Data Orchestration

MORE WEBINARS

Trending Sources

Introducing Menu Maker: Uber Eats’ New Menu Management Tool

Uber Engineering

NOVEMBER 19, 2019

A restaurant’s menu is arguably its most important feature. When ordering online or via the app with Uber Eats, potential customers can’t peer in through a restaurant’s windows or smell the scents wafting from their kitchens, so digital menus become … The post Introducing Menu Maker: Uber Eats’ New Menu Management Tool appeared first on Uber Engineering Blog.

Management

Management Engineering IT Architecture

Webinars

Agent Tooling: Connecting AI to Your Tools, Systems & Data

How to Modernize Manufacturing Without Losing Control

Mastering Apache Airflow® 3.0: What’s New (and What’s Next) for Data Orchestration

MORE WEBINARS

Escaping Analysis Paralysis For Your Data Platform With Data Virtualization

Data Engineering Podcast

NOVEMBER 18, 2019

Summary With the constant evolution of technology for data management it can seem impossible to make an informed decision about whether to build a data warehouse, or a data lake, or just leave your data wherever it currently rests. What’s worse is that any time you have to migrate to a new architecture, all of your analytical code has to change too.

Data Lake

Data Lake Scala Data Warehouse Hadoop

A Guide to Debugging Apache Airflow® DAGs

In Airflow, DAGs (your data pipelines) support nearly every use case. As these workflows grow in complexity and scale, efficiently identifying and resolving issues becomes a critical skill for every data engineer. This is a comprehensive guide with best practices and examples to debugging Airflow DAGs. You’ll learn how to: Create a standardized process for debugging to quickly diagnose errors in your DAGs Identify common issues with DAGs, tasks, and connections Distinguish between Airflow-relate

Data Pipeline

Automated Machine Learning Project Implementation Complexities

KDnuggets

NOVEMBER 22, 2019

To demonstrate the implementation complexity differences along the AutoML highway, let's have a look at how 3 specific software projects approach the implementation of just such an AutoML "solution," namely Keras Tuner, AutoKeras, and automl-gs.

Machine Learning

Machine Learning Project Python

Kafka Streams and ksqlDB Compared – How to Choose

Confluent

NOVEMBER 21, 2019

ksqlDB is a new kind of database purpose-built for stream processing apps, allowing users to build stream processing applications against data in Apache Kafka® and enhancing developer productivity. ksqlDB simplifies […].

Kafka

Kafka Database Process Building

Customer Data Platforms: Silo Killer or Yet Another Silo?

Teradata

NOVEMBER 19, 2019

How do you ensure your Customer Data Platform is enabling breakthrough customer experience business outcomes, rather than hindering them? Find out more!

Data

More Trending

Customer Data Platforms: Silo Killer or Yet Another Silo?

Teradata

NOVEMBER 19, 2019

How do you ensure your Customer Data Platform is enabling breakthrough customer experience business outcomes, rather than hindering them? Find out more!

Data

Netflix at AWS re:Invent 2019

Netflix Tech

NOVEMBER 22, 2019

by Shefali Vyas Dalal AWS re:Invent is a couple weeks away and our engineers & leaders are thrilled to be in attendance yet again this year! Please stop by our “Living Room” for an opportunity to connect or reconnect with Netflixers. We’ve compiled our speaking events below so you know what we’ve been working on. We look forward to seeing you there!

AWS

AWS Entertainment Cloud Software Engineering

Text Encoding: A Review

KDnuggets

NOVEMBER 22, 2019

We will focus here exactly on that part of the analysis that transforms words into numbers and texts into number vectors: text encoding.

Data

Using Confluent Platform to Complete a Massive Cloud Provider Migration and Handle Half a Million Events Per Second

Confluent

NOVEMBER 19, 2019

In the past 12 months, games and other forms of content made with the Unity platform were installed 33 billion times reaching 3 billion devices worldwide. Apart from our real-time […].

Cloud

Cloud AWS

Is There a Geographic Component in Your Analytic Cloud Architecture?

Teradata

NOVEMBER 17, 2019

Moving part of your analytic ecosystem to the cloud requires the inspection of all the ecosystem elements to make sure they perform well over a WAN. Read more.

Cloud

Cloud Architecture

Mastering Apache Airflow® 3.0: What’s New (and What’s Next) for Data Orchestration

Speaker: Tamara Fingerlin, Developer Advocate

Apache Airflow® 3.0, the most anticipated Airflow release yet, officially launched this April. As the de facto standard for data orchestration, Airflow is trusted by over 77,000 organizations to power everything from advanced analytics to production AI and MLOps. With the 3.0 release, the top-requested features from the community were delivered, including a revamped UI for easier navigation, stronger security, and greater flexibility to run tasks anywhere at any time.

Data

Netflix at AWS re:Invent 2019

Netflix Tech

NOVEMBER 22, 2019

AWS

AWS Entertainment Software Engineering Software Engineer

The Math Behind Bayes

KDnuggets

NOVEMBER 19, 2019

This post will be dedicated to explaining the maths behind Bayes Theorem, when its application makes sense, and its differences with Maximum Likelihood.

Geocoding Automation: Free and Paid with Python, Selenium and Google

KDnuggets

NOVEMBER 21, 2019

This tutorial will take you through two options that have automated the geocoding process for the user using Python, Selenium and Google Geocoding API.

Python

Python Process

Three Methods of Data Pre-Processing for Text Classification

KDnuggets

NOVEMBER 21, 2019

This blog shows how text data representations can be used to build a classifier to predict a developer’s deep learning framework of choice based on the code that they wrote, via examples of TensorFlow and PyTorch projects.

Process

Process Deep Learning Data Coding

Agent Tooling: Connecting AI to Your Tools, Systems & Data

Speaker: Alex Salazar, CEO & Co-Founder @ Arcade | Nate Barbettini, Founding Engineer @ Arcade | Tony Karrer, Founder & CTO @ Aggregage

There’s a lot of noise surrounding the ability of AI agents to connect to your tools, systems and data. But building an AI application into a reliable, secure workflow agent isn’t as simple as plugging in an API. As an engineering leader, it can be challenging to make sense of this evolving landscape, but agent tooling provides such high value that it’s critical we figure out how to move forward.

Systems

Data Science for Managers: Programming Languages

KDnuggets

NOVEMBER 19, 2019

In this article, we are going to talk about popular languages for Data Science and briefly describe each of them.

Programming Language

Programming Language Data Science Programming Management

Python Tuples and Tuple Methods

KDnuggets

NOVEMBER 21, 2019

Brush up on your Python basics with this post on creating, using, and manipulating tuples.

Python

Python Programming

The Notebook Anti-Pattern

KDnuggets

NOVEMBER 21, 2019

This article aims to explain why this drive towards the use of notebooks in production is an anti pattern, giving some suggestions along the way.

Python

Generalization in Neural Networks

KDnuggets

NOVEMBER 18, 2019

When training a neural network in deep learning, its performance on processing new data is key. Improving the model's ability to generalize relies on preventing overfitting using these important methods.

Deep Learning

Deep Learning Process IT Data

How to Modernize Manufacturing Without Losing Control

Speaker: Andrew Skoog, Founder of MachinistX & President of Hexis Representatives

Manufacturing is evolving, and the right technology can empower—not replace—your workforce. Smart automation and AI-driven software are revolutionizing decision-making, optimizing processes, and improving efficiency. But how do you implement these tools with confidence and ensure they complement human expertise rather than override it? Join industry expert Andrew Skoog as he explores how manufacturers can leverage automation to enhance operations, streamline workflows, and make smarter, data-dri

Manufacturing

Top KDnuggets tweets, Nov 13-19: A whole lot of Data Science Cheatsheets

KDnuggets

NOVEMBER 21, 2019

Also: Bring the scientific rigor of reproducibility to your Data Science projects; Neutrinos Lead to Unexpected Discovery in Basic Math ; The media gets really excited about AI. Maybe a bit too excited.

Data Science

Data Science Media Data Project

Reproducibility, Replicability, and Data Science

KDnuggets

NOVEMBER 19, 2019

As cornerstones of scientific processes, reproducibility and replicability ensure results can be verified and trusted. These two concepts are also crucial in data science, and as a data scientist, you must follow the same rigor and standards in your projects.

Data Science

Data Science Data Project Process

Neural Networks 201: All About Autoencoders

KDnuggets

NOVEMBER 21, 2019

Autoencoders can be a very powerful tool for leveraging unlabeled data to solve a variety of problem, such as learning a "feature extractor" that helps build powerful classifiers, finding anomalies, or doing a Missing Value Imputation.

Building

Building Machine Learning Data

Deep Learning for Image Classification with Less Data

KDnuggets

NOVEMBER 20, 2019

In this blog I will be demonstrating how deep learning can be applied even if we don’t have enough data.

Deep Learning

Deep Learning Data

The Ultimate Guide to Apache Airflow DAGS

With Airflow being the open-source standard for workflow orchestration, knowing how to write Airflow DAGs has become an essential skill for every data engineer. This eBook provides a comprehensive overview of DAG writing features with plenty of example code. You’ll learn how to: Understand the building blocks DAGs, combine them in complex pipelines, and schedule your DAG to run exactly when you want it to Write DAGs that adapt to your data at runtime and set up alerts and notifications Scale you

Data Engineering

Pro Tips: How to deal with Class Imbalance and Missing Labels

KDnuggets

NOVEMBER 20, 2019

Your spectacularly-performing machine learning model could be subject to the common culprits of class imbalance and missing labels. Learn how to handle these challenges with techniques that remain open areas of new research for addressing real-world machine learning problems.

Machine Learning

Machine Learning Data Preparation Data

Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead

KDnuggets

NOVEMBER 20, 2019

The two main takeaways from this paper: firstly, a sharpening of my understanding of the difference between explainability and interpretability, and why the former may be problematic; and secondly some great pointers to techniques for creating truly interpretable models.

Machine Learning

The Semiconductor Imperative for Driving Meaningful Innovation

KDnuggets

NOVEMBER 20, 2019

The fundamental fact is that more information than ever will need to be analyzed on millions of devices. And that’s where 5G will make accessing data dramatically faster and more efficient. At Samsung, we’re excited about what 5G can truly enable and to be a central player in the new 5G world.

Accessibility

Accessibility Accessible Data

GitHub Repo Raider and the Automation of Machine Learning

KDnuggets

NOVEMBER 18, 2019

Since X never, ever marks the spot, this article raids the GitHub repos in search of quality automated machine learning resources. Read on for projects and papers to help understand and implement AutoML.

Machine Learning

Machine Learning Project Python

Apache Airflow® Best Practices: DAG Writing

Speaker: Tamara Fingerlin, Developer Advocate

In this new webinar, Tamara Fingerlin, Developer Advocate, will walk you through many Airflow best practices and advanced features that can help you make your pipelines more manageable, adaptive, and robust. She'll focus on how to write best-in-class Airflow DAGs using the latest Airflow features like dynamic task mapping and data-driven scheduling!

Data

Why write for KDnuggets? Calling for original blogs and new authors

KDnuggets

NOVEMBER 19, 2019

KDnuggets is calling for original blogs and contributions from new authors on AI, Data Science, Machine Learning, and related topics. The authors of most popular such blogs in December will be profiled in KDnuggets.

Machine Learning

Machine Learning Data Science Data

KDnuggets™ News 19:n44, Nov 20: How I Got Better at Machine Learning; Tips for a cost-effective ML project

KDnuggets

NOVEMBER 20, 2019

Read tips and tricks that helped one Data Scientist to get better at Machine Learning; Learn how to make ML project cost-effective; Consider submitting a blog to KDnuggets - you can be profiled here; and study how to manipulate Python lists.

Machine Learning

Machine Learning Project Python Data

The Reinforcement-Learning Methods that Allow AlphaStar to Outcompete Almost All Human Players at StarCraft II

KDnuggets

NOVEMBER 18, 2019

The new AlphaStar achieved Grandmaster level at StarCraft II overcoming some of the limitations of the previous version. How did it do it?

How to apply machine learning and deep learning methods to audio analysis

KDnuggets

NOVEMBER 19, 2019

Find out how data scientists and AI practitioners can use a machine learning experimentation platform like Comet.ml to apply machine learning and deep learning to methods in the domain of audio analysis.

Deep Learning

Deep Learning Machine Learning Data

How to Achieve High-Accuracy Results When Using LLMs

Speaker: Ben Epstein, Stealth Founder & CTO | Tony Karrer, Founder & CTO, Aggregage

When tasked with building a fundamentally new product line with deeper insights than previously achievable for a high-value client, Ben Epstein and his team faced a significant challenge: how to harness LLMs to produce consistent, high-accuracy outputs at scale. In this new session, Ben will share how he and his team engineered a system (based on proven software engineering approaches) that employs reproducible test variations (via temperature 0 and fixed seeds), and enables non-LLM evaluation m

Software Engineer

Sat.Nov 16, 2019 - Fri.Nov 22, 2019

Advice for New and Junior Data Scientists

Introducing ksqlDB

Webinars

Trending Sources

Introducing Menu Maker: Uber Eats’ New Menu Management Tool

Webinars

Escaping Analysis Paralysis For Your Data Platform With Data Virtualization

A Guide to Debugging Apache Airflow® DAGs

Automated Machine Learning Project Implementation Complexities

Kafka Streams and ksqlDB Compared – How to Choose

Customer Data Platforms: Silo Killer or Yet Another Silo?

Sign up to get articles personalized to your interests!

More Trending

Customer Data Platforms: Silo Killer or Yet Another Silo?

Netflix at AWS re:Invent 2019

Text Encoding: A Review

Using Confluent Platform to Complete a Massive Cloud Provider Migration and Handle Half a Million Events Per Second

Is There a Geographic Component in Your Analytic Cloud Architecture?

Mastering Apache Airflow® 3.0: What’s New (and What’s Next) for Data Orchestration

Netflix at AWS re:Invent 2019

The Math Behind Bayes

Geocoding Automation: Free and Paid with Python, Selenium and Google

Three Methods of Data Pre-Processing for Text Classification

Agent Tooling: Connecting AI to Your Tools, Systems & Data

Data Science for Managers: Programming Languages

Python Tuples and Tuple Methods

The Notebook Anti-Pattern

Generalization in Neural Networks

How to Modernize Manufacturing Without Losing Control

Top KDnuggets tweets, Nov 13-19: A whole lot of Data Science Cheatsheets

Reproducibility, Replicability, and Data Science

Neural Networks 201: All About Autoencoders

Deep Learning for Image Classification with Less Data

The Ultimate Guide to Apache Airflow DAGS

Pro Tips: How to deal with Class Imbalance and Missing Labels

Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead

The Semiconductor Imperative for Driving Meaningful Innovation

GitHub Repo Raider and the Automation of Machine Learning

Apache Airflow® Best Practices: DAG Writing

Why write for KDnuggets? Calling for original blogs and new authors

KDnuggets™ News 19:n44, Nov 20: How I Got Better at Machine Learning; Tips for a cost-effective ML project

The Reinforcement-Learning Methods that Allow AlphaStar to Outcompete Almost All Human Players at StarCraft II

How to apply machine learning and deep learning methods to audio analysis

How to Achieve High-Accuracy Results When Using LLMs

Stay Connected