StreamSets Transformer Released
September 09, 2019
Share this

StreamSets released StreamSets Transformer, a simple-to-use, drag-and-drop UI tool to create native Apache Spark applications.

Designed for a wide range of users — even those without specialized skills — StreamSets Transformer enables the creation of pipelines for performing ETL, stream processing and machine-learning operations. Now, data engineers, scientists, architects and operators gain deep visibility into the execution of Apache Spark while broadening usage across the business.

Apache Spark delivers on the promise of advanced data processing and machine learning at scale. But there are drawbacks. Developing and operating applications on Apache Spark is complex and requires hand-coding. It is typically restricted to developers and companies with mature data engineering and data science practices. In addition, users often have very limited visibility into how their Apache Spark jobs are running. StreamSets Transformer solves these issues. Its easy-to-use, logical user interface and rich tools for designing data transformations eliminate the complexity and need for specialized skills. Pipelines instrumented with StreamSets Transformer provide unparalleled visibility into every Spark execution. Equally important, developers now have a single tool to build both batch and streaming pipelines.

The key features of StreamSets Transformer include:

- Continuous monitoring — Unparalleled visibility into Apache Spark application execution

- Continuous data — Runs in both batch and streaming modes

- Progressive error handling — Finds where and why errors occur without the need for Apache Spark skills to decipher complex log files

- Execute on Apache Spark anywhere — Works in the cloud, Kubernetes or on premises

- Highly extensible — Higher order transformation primitives for the ETL developer, SparkSQL for the analyst, PySpark for the data scientist, and custom Java/Scala processors for the Apache Spark developer

- Sets-based processing — For ETL, machine learning and complex event processing

“With StreamSets Transformer, Apache Spark is finally available to a wide range of users, enabling visibility, monitoring and reporting for mission-critical workloads,” said Arvind Prabhakar, CTO of StreamSets. “In essence, StreamSets Transformer brings the power of Apache Spark to businesses, while eliminating its complexity and guesswork.”

“With StreamSets Transformer and Databricks integrated together, even more users can easily access the powerful capabilities of Delta Lake and our optimized Apache Spark for data science and analytics,” said Michael Hoff, SVP of Business Development and Partners at Databricks. “Especially as organizations migrate from legacy on premises platforms, our partnership will help them efficiently make that transition to manage their data and machine learning workloads in the cloud.”

StreamSets Transformer is available immediately.

Share this

The Latest

April 25, 2024

The use of hybrid multicloud models is forecasted to double over the next one to three years as IT decision makers are facing new pressures to modernize IT infrastructures because of drivers like AI, security, and sustainability, according to the Enterprise Cloud Index (ECI) report from Nutanix ...

April 24, 2024

Over the last 20 years Digital Employee Experience has become a necessity for companies committed to digital transformation and improving IT experiences. In fact, by 2025, more than 50% of IT organizations will use digital employee experience to prioritize and measure digital initiative success ...

April 23, 2024

While most companies are now deploying cloud-based technologies, the 2024 Secure Cloud Networking Field Report from Aviatrix found that there is a silent struggle to maximize value from those investments. Many of the challenges organizations have faced over the past several years have evolved, but continue today ...

April 22, 2024

In our latest research, Cisco's The App Attention Index 2023: Beware the Application Generation, 62% of consumers report their expectations for digital experiences are far higher than they were two years ago, and 64% state they are less forgiving of poor digital services than they were just 12 months ago ...

April 19, 2024

In MEAN TIME TO INSIGHT Episode 5, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses the network source of truth ...

April 18, 2024

A vast majority (89%) of organizations have rapidly expanded their technology in the past few years and three quarters (76%) say it's brought with it increased "chaos" that they have to manage, according to Situation Report 2024: Managing Technology Chaos from Software AG ...

April 17, 2024

In 2024 the number one challenge facing IT teams is a lack of skilled workers, and many are turning to automation as an answer, according to IT Trends: 2024 Industry Report ...

April 16, 2024

Organizations are continuing to embrace multicloud environments and cloud-native architectures to enable rapid transformation and deliver secure innovation. However, despite the speed, scale, and agility enabled by these modern cloud ecosystems, organizations are struggling to manage the explosion of data they create, according to The state of observability 2024: Overcoming complexity through AI-driven analytics and automation strategies, a report from Dynatrace ...

April 15, 2024

Organizations recognize the value of observability, but only 10% of them are actually practicing full observability of their applications and infrastructure. This is among the key findings from the recently completed Logz.io 2024 Observability Pulse Survey and Report ...

April 11, 2024

Businesses must adopt a comprehensive Internet Performance Monitoring (IPM) strategy, says Enterprise Management Associates (EMA), a leading IT analyst research firm. This strategy is crucial to bridge the significant observability gap within today's complex IT infrastructures. The recommendation is particularly timely, given that 99% of enterprises are expanding their use of the Internet as a primary connectivity conduit while facing challenges due to the inefficiency of multiple, disjointed monitoring tools, according to Modern Enterprises Must Boost Observability with Internet Performance Monitoring, a new report from EMA and Catchpoint ...