Advanced Text Processing with Spark - Machine Learning with Spark

Database Reference

In-Depth Information

Chapter 9. Advanced Text Processing with

Spark

In Chapter 3 , Obtaining, Processing, and Preparing Data with Spark , we covered various

topics related to feature extraction and data processing, including the basics of extracting

features from text data. In this chapter, we will introduce more advanced text processing

techniques available in MLlib to work with large-scale text datasets.

In this chapter, we will:

• Work through detailed examples that illustrate data processing, feature extraction,

and the modeling pipeline, as they relate to text data

• Evaluate the similarity between two documents based on the words in the docu-

ments

• Use the extracted text features as inputs for a classification model

• Cover a recent development in natural language processing to model words them-

selves as vectors and illustrate the use of Spark's Word2Vec model to evaluate the

similarity between two words, based on their meaning

Search WWH ::

Custom Search

Home