StreamSets Transformer Released
September 09, 2019
Share this

StreamSets released StreamSets Transformer, a simple-to-use, drag-and-drop UI tool to create native Apache Spark applications.

Designed for a wide range of users — even those without specialized skills — StreamSets Transformer enables the creation of pipelines for performing ETL, stream processing and machine-learning operations. Now, data engineers, scientists, architects and operators gain deep visibility into the execution of Apache Spark while broadening usage across the business.

Apache Spark delivers on the promise of advanced data processing and machine learning at scale. But there are drawbacks. Developing and operating applications on Apache Spark is complex and requires hand-coding. It is typically restricted to developers and companies with mature data engineering and data science practices. In addition, users often have very limited visibility into how their Apache Spark jobs are running. StreamSets Transformer solves these issues. Its easy-to-use, logical user interface and rich tools for designing data transformations eliminate the complexity and need for specialized skills. Pipelines instrumented with StreamSets Transformer provide unparalleled visibility into every Spark execution. Equally important, developers now have a single tool to build both batch and streaming pipelines.

The key features of StreamSets Transformer include:

- Continuous monitoring — Unparalleled visibility into Apache Spark application execution

- Continuous data — Runs in both batch and streaming modes

- Progressive error handling — Finds where and why errors occur without the need for Apache Spark skills to decipher complex log files

- Execute on Apache Spark anywhere — Works in the cloud, Kubernetes or on premises

- Highly extensible — Higher order transformation primitives for the ETL developer, SparkSQL for the analyst, PySpark for the data scientist, and custom Java/Scala processors for the Apache Spark developer

- Sets-based processing — For ETL, machine learning and complex event processing

“With StreamSets Transformer, Apache Spark is finally available to a wide range of users, enabling visibility, monitoring and reporting for mission-critical workloads,” said Arvind Prabhakar, CTO of StreamSets. “In essence, StreamSets Transformer brings the power of Apache Spark to businesses, while eliminating its complexity and guesswork.”

“With StreamSets Transformer and Databricks integrated together, even more users can easily access the powerful capabilities of Delta Lake and our optimized Apache Spark for data science and analytics,” said Michael Hoff, SVP of Business Development and Partners at Databricks. “Especially as organizations migrate from legacy on premises platforms, our partnership will help them efficiently make that transition to manage their data and machine learning workloads in the cloud.”

StreamSets Transformer is available immediately.

Share this

The Latest

November 07, 2019

Microservices have become the go-to architectural standard in modern distributed systems. While there are plenty of tools and techniques to architect, manage, and automate the deployment of such distributed systems, issues during troubleshooting still happen at the individual service level, thereby prolonging the time taken to resolve an outage ...

November 06, 2019

A recent APMdigest blog by Jean Tunis provided an excellent background on Application Performance Monitoring (APM) and what it does. A further topic that I wanted to touch on though is the need for good quality data. If you are to get the most out of your APM solution possible, you will need to feed it with the best quality data ...

November 05, 2019

Humans and manual processes can no longer keep pace with network innovation, evolution, complexity, and change. That's why we're hearing more about self-driving networks, self-healing networks, intent-based networking, and other concepts. These approaches collectively belong to a growing focus area called AIOps, which aims to apply automation, AI and ML to support modern network operations ...

November 04, 2019

IT outages happen to companies across the globe, regardless of location, annual revenue or size. Even the most mammoth companies are at risk of downtime. Increasingly over the past few years, high-profile IT outages — defined as when the services or systems a business provides suddenly become unavailable — have ended up splashed across national news headlines ...

October 31, 2019

APM tools are ideal for an application owner or a line of business owner to track the performance of their key applications. But these tools have broader applicability to different stakeholders in an organization. In this blog, we will review the teams and functional departments that can make use of an APM tool and how they could put it to work ...

October 30, 2019

Enterprises depending exclusively on legacy monitoring tools are falling behind in business agility and operational efficiency, according to a new study, Prevalence of Legacy Tools Paralyzes Enterprises' Ability to Innovate conducted by Forrester Consulting ...

October 29, 2019

Hyperconverged infrastructure is sometimes referred to as a "data center in a box" because, after the initial cabling and minimal networking configuration, it has all of the features and functionality of the traditional 3-2-1 virtualization architecture (except that single point of failure) ...

October 28, 2019

Hyperconvergence is a term that is gaining rapid interest across the manufacturing industry due to the undeniable benefits it has delivered to IT professionals seeking to modernize their data center, or as is a popular buzzword today ― "transform." Today, in particular, the manufacturing industry is looking to hyperconvergence for the potential benefits it can provide to its emerging and growing use of IoT and its growing need for edge computing systems ...

October 24, 2019

More than 92 percent of US respondents agree that Artificial Intelligence (AI) and Machine Learning (ML) will become important for how they run their digital systems ...

October 23, 2019

Progress has been made with digital transformation projects, however technology leaders are finding that running their digitally transformed organizations is challenging and they are under increased pressure to prove business value, according to a survey from New Relic ...