CIS 493/593

Big Data (3-0-3)

Course Content

  • Class Announcement and Post
  • Class Syllabus
  • Group Project
  • Lab Assignments
  • Class Lecture Notes


  • Class Announcement and POST



    IMPORTANT NOTE !!!

    If You Have a 404 Error in any URL in the Class Lecture Notes under http://eecs.csuohio.edu/~sschung/, You have to replace "http" with "https" to be able to access !


    10. April 25, 2023

    The Final Exam on Monday May 8 at 4:00PM - 6:00PM

    In Each Lecture Note for the Final, Only the Slides (Subjects) Covered in Class Will Be in the Final

    See the Links for Final Exam Lecture Note Only (Scroll Down to the Later Half of Lecture Note Sections)


    Topics to Focus on for the Final of CIS493/593

    Design of Intelligent System with data pipelining of Big data processing
    Unstructured Text Processing Techniques, Data Preprocessing Methods
    Inverted Index
    TF-IDF, Cosine Similarity
    NLP Text Analysis methods - POS Tagging
    Data Preprocessing Methods for Classification with ML Algorithms
    Classification with Machine Learning Algorithms: Decision Tree and Neural Network
    Map Reduce Process on Hadoop Distributed File System

    One Page Note is Allowed to the Final



    10. Jan 16, 2023

    Information on Midterm


    Midterm will Be on March 8 at 4:00PM - 5:30PM !

    See Lecture Links Only for MidTerm (Scroll down to the Lecture Note Section to See Active links - All the Rest of Links are Disabled)


    The Topics to Focus On for the Midterm:

    Basically whichever Lecture notes (Slides) covered in class!

    Characteristics of Big Data and their Contents, Main Differences between Big Data and Traditional Data, Common Big Data Applications/Architecture,
    Universal Data Exchange Formats in Object Exchange Model(OEM): Three Common Big Data formats in Semi-Structured Model:
    - HTML, XML, JSON - Data Model, Syntax of Each Encoding Format, all the related data processing techniques -- DOM, XPath.
    Characteristics of Semi-Structured Data Model, Main Differences between Semi-structed Data Model and Relational Data Model
    for Object Relation Mapping (ORM), Conversion between Relation(CSV), XML, and JSON

    Unstructured Data Processing, Inverted Index, How to Build Inverted Index, Text Cleaning Preprocessing, Basic Natural Language Processing (NLP) Methods in Data Pipeling, POS Tagging
    Comparison Between Semistructured Database Server and Relational Database Server for Database Management Strategies.
    Semi-Structured Database System -- MongoDB: CRUD Basic Operations, MongoDB Queries for Embedded Objects and Array, Aggregation Pipelining

    NOTE that ONE Page Note is NOT allowed to the Midterm !!



    Exam Format:

    6-7 Main Questions with 2-3 subquestions for a Small/Short Answer Types for Probelm Solving



    8. Jan 16, 2023

    Only the registered students can access the course blackboard.
    If you have a problem with your blackboard, please contact the registrar or e-learning center CSU Tech Support to resolve the issue !

    Faculty does not control your registration and the course blackboard access in the CSU systems.

    e-learning CSU Tech Support


    TA info:

    9. Jan 16, 2023:

    TA Information

    TA: Yixi Luo

    Email: luoyixi.cn@gmail.com

    Office Hours: Mon, Wed 9:50 AM - 11:50 PM Tentatively (Send email to TA ahead to set up a time slot to let him/her know that you are coming)

    Location: Big Data Analytics Lab: FH 305 or Recommend a ZOOM Meeting for Your Safety !

    ZOOM Meeting ID: 306 371 3308
    Password: VFw6s8
    ZOOM Meeting Link

    If you have questions in Labs or grading your Labs, Send an email to TA to See During the TA's Office Hours or Schedule a Zoom meeting



    Dr. Chung's Office Hours:

    Mon and Wed 1:30PM - 3:30PM

    Zoom Meeting Only for Safety ! Send Me Email to Set Up a Zoom Meeting.
    Meeting ID: 859 3867 6332

    Email: s.chung@csuohio.edu


    6. Jan 16, 2023:
    Lab Submission:

    The Output of each lab is your report in Doc file that shows your screen captures of each of your executions with Your Outputs in your web browser or the server returned the correct results in the webpage.
    your Report in Doc file also should expain all the platform set up, the execution steps, and copy of each source code files ) Each of your screen capture must show your results returned in your web browser of your system to prove that your lab is done correctly !!

    Submit your Zip file that includes the followings on Blackboard for a timestamp and as a proof
    1) Your Lab Report in .doc file that explains all the platform set up procedures, the execution steps, each intermediate output, final outputs, and a copy of each source code files,
    2) All your Source files, and output files


    5. Jan 16, 2023:
    See Updated Instructions for File Permission Error 403 in the Lab Section!

    4. Jan 16, 2023:
    Instructions to Set Up Webpage is Posted in the Lab Section!

    3. Jan 16, 2023:
    Each Class Attendance is Required for CIS492/593 !
    The ZOOM Meeting Log for Each Class Shows Your Class Attedance
    There Will Be Random Quizzes or Sign Up Sheet to Check the Attendance As Well !

    2. Jan 16, 2023:
    You can use any SQL Server for your Database such as MySql or MS SQL Server. See MySql Set up Guide in the Lab section.

    1. Jan 16, 2023:
    The class webpage for this semester will be announced on blackboard:
    CIS492/593 Big Data Class Webpage

    Or You can Reach from Teaching Section of My Website for Big data Research Lab at:
    Big Data Research Lab

    If you have a trouble to display the webpage correctly with MS Internet Explorer, open it with Google Chrome

    Check the Last Day to Add and Drop here !
    University's Official Academic Calendar for the Semester and the Final Exam schedules

    Final Exam Schedule:
    Mon May 10, 4:00 - 6:00 PM




    Information on Midterm and Final Exams


    7. Midterm will Be Tentatively on March 8 !

    The Topics to Focus On for the Midterm:

    Basically whichever Lecture notes (Slides) covered in class!

    Characteristics of Big Data and their Contents, Main Differences between Big data and Traditional Data, Common Big Data Applications/ architecture,
    Universal Data Exchange Formats in Object Exchange Model(OEM): Three common big data formats as semi structured data:
    - HTML, XML, JSON - Data Model, Syntax of Each Encoding Format, all the related processing techniques -- DOM, XPath.
    Characteristics of Semi-Structured Data Model, Main Differences between Semi-structed Data Model and Relational Data Model
    Comparison between Relational Data model and Semi-Structured Data Model. Conversion between Relation(CSV), XML, and JSON

    Unstructured Data Processing, Inverted Index, Text Cleaning, Basic NLP Processing Methods in Data Pipeling, POS Tagging
    Comparison Between Semistructured Database Server and Relational Database Server for Database Management Strategies.
    Semi-Structured Database System -- MongoDB: CRUD Basic Operations, MongoDB Queries for Embedded Objects and Array, Aggregation Pipelining

    NOTE that ONE Page Note is NOT allowed to the Midterm !!




    See the Links for Final Exam Lecture Note Only


    Topics to Focus on for the Final of CIS493/593

    Design of Intelligent System with data pipelining of Big data processing
    Unstructured Text Processing Techniques, Data Preprocessing Methods
    Inverted Index
    TF-IDF, Cosine Similarity
    NLP Text Analysis methods - POS Tagging
    Data Preprocessing Methods for Classification with ML Algorithms
    Classification with Machine Learning Algorithms: Decision Tree and Neural Network
    Map Reduce Process on Hadoop Distributed File System

    One Page Hand Wrtten Note (Each Side) is Allowed to the Final !
    Printed Copy of Lecture Notes Are NOT Permitted as One Page Note





  • (The course number will be changed to CIS468/568 from Spring 2024 as a Regular Course)
  • Syllabus of CIS492/593 Big Data
  • ABET Syllabus of CIS/DSA 468 Big Data (This course will be offered as DSA468 from Spring 2024 as a core course of BS Data Science Degree)

  • Project and Presentation



    Projects on Big Data Processing, Building an Intelligent Web Application with Big Data Analytics


    The Scope of Group Project Requirements Has Been Adjusted to: (This Adjustment May not Apply to this semester. Will Be Announced)

    1. You Can Do One Person Group Project in a Small Scale.
    2. a Project with a Small Scale with Some Extension of One of Lab4 - Lab4_3
    3. For Those Who Have Aleady Taken CIS660, it is Required to Complete a Full Project with Bid data set and Presentation
    4. Implementing Your Final Project with Clientside and Serverside of a Web Application with a Web Based User Interface Is NOT Required. It will Be Counted as Extra Credit.


    Final Group Project Specification and Instructions:

    Project Submission Instructions:
    Submit Group Project Presentation and Final Report in a Zip File By the End of Friday of Your Presentation Week !
    Remember you have to include the source file of your Project Report in doc and Presentation slides in pptx !
    If your data file is too big to upload, Submit your zip file with your Data file either on the google drive or One Drive and send the link. (Send me email for access permission for this link !)

    Submit a Zip file that includes:
    All of your presentation slides (both in .ppt and .pdf) and
    Your Group Final Project Report (in doc) with Platform/System Set up Procedures/Instructions, Executions Steps, all the source codes, scripts, all the intermediate outputs, and final output files on Blackboard by the end of Friday of your presentation week.
    Include the Problems/Error Encountered and Your Resolutions in Your Report
    One Submission Per Group Required.
    Submit a Zip File that Includes All the required Source Files, Input, Output Files, and Final Report (in Doc file) and Presentation Slides (in pptx).
    Your Final Project Report Should Include the Set Up Procedure /Configuration Detail of Your Platform/System/Packages as well as Source Codes and Intermediate Results in files.
    The Report Should Explain Each Step of Your Project Tasks with the Screen Captures and Results.
    If you don't show/include any of the required contents in your report and presentation, I will ASSUME that your group submitted a Copy of Somebody's Github Codes your group downloaded from the Web.

  • Project Specification and Instructions on What To Submit
  • Project Task Description and Suggested Project List (Will be Added More)
  • FAQs for Final Project

  • Important Notes for Final Project:

    - 1 - 3 Person Group Project Are Allowed
    - One Person Final Project is Allowed. You Can Work Alone.
    - Three Person Group is Allowed. Make Sure to Make the Project in a Bigger Scale
    - You can Change Your Project Plan/Proposal or the Details even after Your Proposal is submitted until the Deadline of the Status Report.
    - If You Need to Find Group Members, Use Blackboard Email to Send to the Class, Some will respond to you if they are looking for a group member.



    The IMDB Data links in the Project List are gone. Check here for Movie Review Data Set.
    Movie Review Data
    Youtube Analytic Site


    For Those Who Have Already Taken CIS660, The Final Group Project Should Include Fully Analytic Processing

    Suggested Projects:

    For Text Analytics like Sentiment Analysis or Opinion Analysis: NLP Techniques - POS, NER Tagging, Bi-Gram Handling Are Required for Preprocessing.
    For Document Categorization: by Constructing TF-IDF Vectorization. Inverted Index Building Will Be Plus but Optional.
    Building Word2Vec Embeddings for a Collection of Documents/Webpages with Training Set Generation in Skip Gram Model
    For Other Types of Projects, the Proposal is Required to Be Approved to Meet the Complexity of Final Project


  • Requirement for Final Project Report (See The Google Presentation Schedule Sheet or the Final Submission Instruction in the Project Section
  • Group Project Presentation Schedule (Invitation to One Drive Sign Up Sheet for Presentation Schedule Sheet Will Be Sent to your CSU CampusNet Email)

  • Extra Credit Final Project:
  • Extra Credit Final Project (Recommended for the Honor/Scholar Contract Course Students)
  • Example of Extra Credit Final Project


  • Final Group Project Time Line: (Tentative)

    Group Project Proposal Due by April 7 !
    Group Project Status Report Due by April 21 !
    Group Project Presentation Either on May 1 or May 3 !
    Final Group Project Report Due By Friday May 5 !




    Any Group Size in 1 - 4 Person Is Allowed
    Since the Size of the Class is Too Big, 3-4 Person Project Group Will Be Allowed As Long As the Project Scale is Big Enough for 3-4 Persons


    IMPORTANT Submissions for Group Project

    Task 1: Group Project Proposal (Plan)
    Submit minimum 3 Page Group Proposal on Blackboard with List of Group Memebers

    Group Project Proposal (Plan) Should Include Brief Descriptions on:
    1. Description of Big Data in Size and Format, Data Collection Plan, 2. Goal of Your Big Data Analytic Project with What Kind of Intelligent Analytic Funtionality, Features of Your AI/Big Data Analytic Application
    3. Big Data Processing Plan, Methods
    4. Investigate on Platform/Systems/Tools/APIs to Use


    Task 2: Group Project Status Report

    Your Project Status Report Should Show the Following Tasks Done:
    1. Platform Setting/System Configuration Procedure (if it is new)
    2.Your KnowledgeBase Structure/ Database Design
    3. Design of Big Data Processing Pipeline, Data Transformation Methods
    4. Your Server Data/Database Contents in Progress, Any Intermediate Outputs in Progress


    Task 3: Project Presentation Project Presentation Starts from the Last Week of the Class of the Semester
    Project Presentation Schedule Will Be Sent To Your CSU Email for Sign Up One week Before the Presentation
    Read the Instructions of Project Presentation Here !


    Group Project Presentation Should Include:

    1. Data Description, Data Size, Data Collection Method
    2. Platform Setting/System Configuration Procedures
    3. System Design (Architecture) of Your AI Application or Data Analytic Goal in Detail
    4. Raw Big Data Preprocessing Methods and Intermediate Results
    5. Design of Big Data Processing Pipeline, Data Transformation Methods 6. Description of Your KnowledgeBase Structure/Database Design. And Show the Contents
    7. Ranking Algorithm, Data Matrix (Structures) if any for Evaluation
    8. The Problems/Errors Encountered and Your Resolutions
    9. System Demo or Evaluation Results and Visualization of the Result




    Submit Final Project Report and Presentation By the End of Friday of The Last Class Week After Your Presentation!

    One Submission Per Group Required.
    Submit a Zip file that includes:
    1. All of your presentation slides (both in .pptx) and
    2. Your Group Final Project Report (in doc) 
    
    The Final Project Report Should Explain Each Step of Your Project Tasks with the Screen Captures and Results.
    
    The Report Should Include Platform/System Set up the Set Up Procedure /Configuration Detail of Your Platform/System/Packages, Executions Steps, all the Source Codes, Scripts, all the intermediate outputs, and final output files.
    Include the Problems/Error Encountered and Your Resolutions in Your Report
    
    If you don't show/include any of the required contents in your report and presentation, I will ASSUME that your group submitted a Copy of Somebody's Github Codes your group downloaded from the Web.
    
    Group Project Presentation Should Include:
    
    1. Data Description, Data Size, Data Collection Method
    2. Platform Setting/System Configuration Procedures
    3. System Design (Architecture) of Your AI Application or Data Analytic Goal in Detail
    4. Raw Big Data Preprocessing Methods and Intermediate Results
    5. Design of Big Data Processing Pipeline, Data Transformation Methods 6. Description of Your KnowledgeBase Structure/Database Design. And Show the Contents
    7. Ranking Algorithm, Data Matrix (Structures) if any for Evaluation
    8. The Problems/Errors Encountered and Your Resolutions
    9. System Demo or Evaluation Results and Visualization of the Result
    






    <
    Requirements of Combining two Group projects of two related courses:

    Combining two Group projects of two related courses (CIS492 with CIS408 or CIS593 Deep Learning) are ok as long as the combined project focuses on both sides of the subjects
    for example, with Web application aspect for CIS408 and Server-side big data processing with a database server and any analytic methods with big data covered in CIS492.

    The data of the Deep Learning class CIS593 is images, which is a lot different from the Big data covered in this Big data class such as a large volumn of semistructured collection or text documents.
    Maybe training a large text documents using deep learning would be a good combined project of CIS492/CIS593 and the Deep Learning Class.




    For Extra Credit Projects or Contract Course Requirement for Honor Students

    It includes a project to build a real-life web based AI applications like:

    - Web Search Engine for a certain domain, for example, .csuohio.edu or .cnn
    - WebMD
    - Question Answering system like Alexa but in a webbased interface to get a user question





    Project Examples of Vectorization of each document in a Training set for a Machine Learning Classifier:

    Project Example I: Sentiment Analysis of Yelp Business Review with Machine Learning
    Project Example II: Sentiment Analysis of Yelp Business Review with Machine Learning


    Best Senior Design Projects on Big Data Processing and Text Anaytics (Created from the Subject of CIS493 and CIS408):


  • 2017 Senior Project: Wikipedia Search Engine (Created From the Final Project of CIS408 and CIS493 Subjects) by Nick McCoy, et al
  • The Candidate of Best Senior Project in Engineering College of 2019 (Created From the Final Project of CIS408 and CIS493 Subjects) by Joel Stell et. al.
  • The Best CS Senior Project Winner of 2017 (Created From the Final Project of CIS408 and CIS493 Subjects) by Mike D'Arcy and Utkarsh Patel
  • The First Prize Winner of 2016 Senior Project From CIS430 and CIS408 by Nick White (Now in FaceBook), et al


  • Best Group Projects: Selected Best Projects Will be Posted Here !

    Big Data and Data Science Projects:

  • Social Media Opinion Analysis System for 2020 Presidential Election Prediction (the Candidate of Best Senior Project in Engineering College of 2021 (Created From the Final Project of CIS408 and CIS430)
  • 2021 Senior Project: Stock Market Analysis Service System (From the Final Project of CIS408 and CIS430)
  • Intelligent Infectious Disease Tracking System (From CIS408 and CIS430)
  • Product Review Sentiment Analysis System (From CIS430 and CIS408)

  • Sharesci: Online Research Document Search Engine in Natural Language (with Mongo DB, Node Js and Angular JS) by Mike D'Arcy and Utkarsh Patel


  • You Can Choose to Extend One of The Extra Credit Labs Below as a Final Group Project either with a Different data set or the same data set

    Extra Credit Lab 4_2 on Webpage Categorization by Topics

  • Lab 4_2 on Cosine Simiarity Measure for Webpage Clustering (Document Categorization) by Topics


  • Extra Credit Lab 4_3 on Information Retrieval Methods for Content Based Document Search Engine


  • Lab 4_3 on Content Based Document Search Engine with Collection of Union Address


  • Example of Inverted Index on State Union Addresses (Note that the Structures of the Index Tables are a little Different in the Example)





    Extra Credt Lab 5 on Sentiment Analysis with Machine Learning for Classification:
  • Extra Lab 5 on Sentiment Analysis with Machine Learning for Classification

  • Data Set: Choose a Review Text Data Set Obtained from Social Network sites: Twitter or Yelp Review Data Set


    Classification Goal:

    1. For each review text obtained from Twitter texts, Derive a Preditive Model to Predict (Classify) Whether it is Positive or Negative Review.
    2. For each review text in Yelp Review Data Set, Derive a Preditive Model to Predict (Classify) the scale of the Review in 1 - 5. (star)




    More To Come Here !


    Group Project Data Sources: You can choose to work on these data sets for your group project

  • Wiki Page Data set
  • Complete Yelp Callenge Data set: 5 BIG JSON Files
  • More to Come Here !


    Final Project Submission Instructions:

    Submit Group Project Presentation and Final Report in a Zip File By the End of Friday of Your Presentation Week !
    Remember you have to include the source file of your Project Report in doc and Presentation slides in pptx !
    If your data file is too big to upload, Submit your zip file with your Data file on your google drive or One Drive and Send email to me and TA to share !

    Submit a Zip file on Blackboard by the end of Friday of your presentation week.
    One Submission Per Group Required.

    Your Project Zip File Should includes:

    1) All of your presentation slides (both in .ppt and .pdf) and
    2) Your Group Final Project Report (in doc) with Platform/System Set up Procedures/Instructions, Executions Steps, all the source codes, scripts, all the intermediate outputs, and final output files
    3)Include the Problems/Error Encountered and Your Resolutions in Your Report

    Your Final Project Report Should Include the Set Up Procedure /Configuration Detail of Your Platform/System/Packages as well as Source Codes and Intermediate Results in files.
    The Report Should Explain Each Step of Your Project Tasks with the Screen Captures and Results.

    IMPORTANT NOTE !!!
    If you don't show/include any of the required contents in your report and presentation, I will ASSUME that your group submitted a Copy of Somebody's Github Codes your group downloaded from the Web.


  • Lab Assignments and Resources



    Basic Python Tutorials :

  • Python Tutorial
  • Python Codecademy Tutorial Site



  • Popular Python Data Science Platforms:

    Installation Guide for Scientific Python tools for Data Science with pip (inbuilt package management system)

    • Anaconda
    Anaconda Open Source Site See Fundamental Section for List of Data Science Platforms
    Anaconda Tutorials

    • Python Anaconda Tutorial Sites
    Anaconda Tutorial Site
    Anaconda Tutorial Site

    Basic Guide for Python tools for Data Science: jupyter-notebook
    More Basic Guide for Python tools for Data Science: jupyter-notebook

    • Python Scikit Learn for Common Data Science Tasks
  • Python Scikit Learn
  • Python Scikit Learn for Data Preprocessing
  • Python Text Processing Libs for Text Analysis
  • Python Numpy Tutorial

  • Text Preprocessing (Natural Language Processing) Library in Python SpaCy:

    Liquistic Modules in Python SpaCy
    Lemmatizer in Python SpaCy
    Liquistic Modules for Tokenization, Stemming, Lemmatization in Python SpaCy
    How to Code Liquistic Modules like Lemmatizer in Python SpaCy
    Python Example for Basic Text Processing

    • Python sklearn.cluster Python Sklearn Clustering



    Python IDE Deduggers:

    Basic Guide for Python Debugger Pycharm
    Python IDE Spyder
    Python Debugger Spyder




    Basic R Tutorials :

  • R Studio basic Tutorial
  • R Basic Online Lecture

  • R Manuals

  • R Tutorial with Examples with R Stat Tool
  • Note that Examples in this tutorials may not the final correct output for Lab1 !
  • Examples of Basic R Stat Tool with Helpful References

  • Special Online Study Guides on Basics on Data Warehouse/OLAP, Data Analytics, Big Data in Independent Study Independent Study with Nick White (Now in FaceBook and The First Prize Winner of 2016 Senior Project)



    Useful Machine Learning Tutorial Sites:
    Keras for Image Processing/Text Processing with Deep Learning:
    Keras Machine Learning


    For Your Own Advanced Study
  • Coursera Machine Learning by Stanford
  • Open AI
  • Google Colab
  • Medium by MIT



  • Lab Submission Instructions:


  • The Output of each lab is your Lab Report in Doc file that shows your screen captures of each of your executions with Your Outputs

  • Your Report in Doc file should include all the platform set up procedures, the execution steps, and copy of each source code files
  • Each of your screen capture must show your results returned by your systems and your database servers to prove that you have done the lab correctly !!



  • 1. Submit your Zip file that includes your report in .doc file (that expains all the platform set up, the execution steps, and copy of each source code files ) and all the Source files, and output files on Blackboard for a timestamp and as a proof.

    2. IMPORTANT NOTE:
    Your Screen Captures in Your Lab Report Should Show Your Own System and Your Database Server Name to Prove That Your Lab Was Done In Your System.

  • Example of Output of the Execution Steps for Labs

  • 3. If You did Extra Credit Part, Mention about What Part is Done for Extra Credit at the Front Page of Your Report in Bigger and Bold Font !



    Useful Lab Helpers:

  • XHTML Validator
  • How to debug HTML, JAVASCRIPT, DOM, XPath
  • JavaScript Debugger
  • Node.JS Setup Recommendation
  • MDN Site for DOM with XPath in JavaScript
  • XPath Setup Guide for JavaScript
  • XPath Methods
  • Setting Up DOM with XPath for Python
  • Webscapping with DOM, XPath for Python with BeautifulSoup
  • Python XPath Guide

  • Useful Tools: Swagger API Tool for REST API Developments

    Useful Big Data Analytic Tools

    Choose your System/Tool/Platform to Set Up and Get Used to:

    Machine Learning in Python:

  • Python Scikit Learn
  • Python Scikit Learn for Data Preprocessing
  • Python Numpy Tutorial

  • Natural Language Processing (NLP) for Text Preprocessing/Big Data Analytics Library in Python SpaCy:

  • Python Text Processing Libs for Text Analysis
  • Liquistic Modules in Python SpaCy
    Lemmatizer in Python SpaCy
    Liquistic Modules for Tokenization, Stemming, Lemmatization in Python SpaCy
    How to Code Liquistic Modules like Lemmatizer in Python SpaCy
    Python Example for Basic Text Processing









    Lab Assignments:


    The Lab Submission Link and the Deadline of Each Lab Will Be Posted on the Class BlackBoard !
    Wait For the Lab Submission Links Are Created on Blackboard for Each Lab









    Lab0: Learning Python -- Due by the End of the Second Friday of the Semester

  • Python Codecademy Tutorial Site

  • Python Data Science Platforms: See More Python Platforms Above or Lab1 Set Up Guides Below to Choose for Data Science

    Installation Guide for Scientific Python tools for Data Science with pip (inbuilt package management system)

  • Installation and Setting Up Either MySQL or MS SQL Server

  • Installation Guides for MySql Server:

    MySql Download
    How to Create MySQL Database


    If You Want to Use MS SQL Server, Installation Guides for MS SQL Server:

    See the Announcement Section of CIS430/530 Database Systems and Processing below for Account Creation for Microsoft Azure site for Free Download of MS Visual Studio and MS SQL Sever.
    How to Create MS Azure Potal Site Account

    See the Lab Section of CIS430/530 for Installation Instruction of MS Visual Studio and MS SQL Sever.
    How to Download and Install MS SQL Server



    If you have an issue with your own computer, you can borrow a Laptop from the university Tech center. Or you can set up on the Azure Cloud or Amazon Cloud.
    There are the general CS computer labs in the Fenn Hall at the first floor (if they are open over the pandemic). However, any computer lab won't allow you to download and set up your own platform with a database server.
    The subjects of Big data are advanced and new, the course requires to set up a new system like MongoDB or Hadoop in your own system.










    IMPORTANT NOTE:
    Your Screen Captures in Your Lab Report Should Show Your Own System and Your Database Server Name to Prove That Your Lab Was Done In Your System.











    If you have an issue with your own computer, you can borrow a Laptop from the university Tech center. Or you can set up on the Azure Cloud or Amazon Cloud.
    There are the general CS computer labs in the Fenn Hall at the first floor (if they are open over the pandemic). However, any computer lab won't allow you to download and set up your own platform with a database server.
    The subjects of Big data are advanced and new, the course requires to set up a new system like MongoDB or Hadoop in your own system.



    IMPORTANT NOTE:
    Your Screen Captures in Your Lab Report Should Show Your Own System and Your Database Server Name to Prove That Your Lab Was Done In Your System.




    The Lab Submission Link and the Deadline of Each Lab Will Be Posted on the Class BlackBoard !
    You Have to Start Working on Labs Before the Submission Link Are Created on Blackboard for Each Lab Submission.

    If You did an Extra Credit Part, Mention about What Part is Done for Extra Credit at the Front(Cover) Page of Your Report in Bigger and Bold Font !

    Always Follow the Deadline of Each Lab Assigned on the Class Blackboard.
    The Deadlines mentioned on the Class Webpage Are Tentatively Scheduled at the Beginning of Each Semester.

    Please Identify Your Course When You Ask Me in Email !












    Lab 1 on Web Data Processing for Information Extraction

  • Lab Assignment 1

  • FAQs for Lab1 on Information Extraction from Webpages

  • The Simplied html file of the Info site (You may use this simplified html file to inspect the element structure to extract the required info for Lab1)

  • Infoplease site of State Union Addresses of US Presidents
    Correct page of Address of John Adams December 3 1799

    Important Notes:
    1. Part 2 (on Combining all the address texts in one text file) Is Required For CIS593 Students.
    Part 2 Is NOT Required for CIS492 Students but It is for extra credit for CIS492 Students.

    2. If there is a Link that Does NOT Have Any Web Page Contents, Add NULL Values for the corresponding Columns for the link

    3. Do not Assume that every sites has an identical URL format. This is semi-structured data. Nothing is regular in Big data.
    For the irregular parts, use regular expressions or xpath as neccessary.


    There are two ways to do Information Extraction from Webpages. Either Method is fine for Lab1 !

    Method 1: Webpage as Semi-Structured HTML DOM Tree Using XPATH -- Extra Credit !!
    Method 2: Webpage as Unstructured Text Using Parsing API like Beautiful Soup


    Python Setup Guides for Labs:

    pip Python package Installation Guide
    Setup Guide for Anaconda Python Framework and Jupyter Notebook IDE (by TA Hemal Paneliya)
    Setup Guide for Anaconda Python Framework and Spider IDE and Debugger (by TA Durga Dasepalli)

    How to Set Up Python Jupyter Notebook to Connect SQL Server using pyodbc



    Python IDE Deduggers:

    Basic Guide for Python Debugger Pycharm
    Python IDE Spyder
    Python Debugger Spyder


    Database Server Set up is needed for the Labs.
    You can use any SQL Server -- MySQL, MS SQL Server, or any Database Server

    See Lab0 Section Above for more instructions or See the step-by-step installation guides in CIS430/530 Lab Section Below
    CIS430/530 Lab Section
    How to Download and Install MS SQL Server



    Note that You don't Need to Do Any Extra Set Up to Use XPATH in a JavaScript in a Client side Codes for HTML DOM Processing which will be executed by your Webbrowser.
    All the Recent Web Browsers Have XPATH Features in their Debugger by Default (Since 2017).

    You Need To Set UP DOM and XPATH for Any Server-side Scripts/Languages for Applications such as Lab1.

    Installation Guide for lxml with pip Python package for DOM with XPath and pycharm IDE for debugging

    Lab1 Implementation Guides:

    Note that the Code Examples of the Lab1 Guides below Do NOT Contain the Complete Codes to Be Executed. These are ONLY for the guides for Lab1.
    The Environment Configuration and the Versions of APIs Varies.
    Do Not Copy the Entire Codes Blindly to Do Your Lab1 since it will not work depending on the version of your python and setting up for DOM and XPath.

    Code Examples with DOM and XPATH:
    Example of Lab1: General Guide in Python with DOM and XPath, and pyodbc for database operations for Information Extraction
    Example of Lab1: Guide in Python with DOM and XPath in lxml for Information Extraction
    Example of Lab1: Guide in PHP with DOM and XPath for Webscapping
    Example of Lab1: Guide in Python on Linux for Webscapping

    Code Examples with Beautiful Soup Text Processing API:
    Example of Lab1: Guide in Python with Beautiful Soup for Information Extraction



    DOM with XPATH Set Up in Any Script/Programming Languages for Server Side Applications
    Note that You don't Need to Do Any Extra Set Up to Use XPATH in your Serverside JavaScript in NodeJS.


    All the labs of CIS492/593 are Server-side data processing in a server-side script/programming language with file I/O and DOM parser with Xpath. They are NOT a clientside Javascript executed by your Webbrowser.
    For the Labs to Write an Application Server in this course, You Need To Add DOM Parser related APIs for Your Scripts/Programming Languages To Do the following steps:
    to Call a DOM Parser API
    to Build a DOM Tree and
    to Use XPATH Methods for Retrival to Extarct information You need in the Application Codes as in the Examples below.

    Any Modern Script/Programming Languages Have the DOM Parser and XPath Features. You Need to Set up to Use.

    Setting Up lxml parser for DOM with XPath for Python
    Node.JS XPath Setup Guide and XML Parser in JavaScript based Node JS
    XPath Setup Guide for JavaScript
    XPath Setup with npm in JavaScript for Node JS


    Set Up DOM with XPath for HTML and XML:

    How to Debug HTML with JavaScript, DOM, XPath
    Learn How to Inspect Element in Chrome to Debug HTML with for DOM and XPath

    JavaScript:
    JavaScript Debugger
    XPath Examples in Javascript
    MDN Site for DOM with XPath in JavaScript
    DOM with XPath Setup Guide for JavaScript
    XPath Methods

    Node.JS XPath Setup Guide
    npm Node.JS XPath Setup Guide
    Google Crome XPath Helper Setup

    Python:
    Setting Up DOM with XPath for Python
    BeautifulSoup Documentation for HTML, XML for Python
    Webscapping with DOM, XPath for Python with BeautifulSoup (From the Stanford Class)
    Python XPath Guide


    Automatic Table Creation in a SQL Server

    After installing a SQL server (See the step-by-step installation guides in CIS430/530 Lab Section
    Add the codes in pyODBC as below for a Connection for the Anaconda Python Framework to Connect to Your SQL Database Server

    For a connection with SQL Server with servername and database name using pyodbc in Python:

    conn = pyodbc.connect('Driver={SQL Server};Server=YOUR_SQL_SERVERNAME\SQLEXPRESS;Database=YOUR_DATABASE_NAME;Trusted_Connection=yes;')

    pyodbc(Open Database Connectivity) to Connect MS SQL Server in Python:

    Installation pyodbc to Connect to MS SQL Server in Python
    pyodbc to Connect to MS SQL Server in Python
    How to Connect to MS SQL Server in Python
    pyodbc: Examples of ODBC in Python

    Column Data Types to store a Large Text data to Create a Table with in MS SQL Server or other database server

    How to Create a Table in a SQL Server from CSV/TSV Text files

    How to Create a Table from a File with Bulk Insert with MS SQL Server
    How to Create a Table from a file with Bulk Insert in MySQL Server
    Example of a Script to Create Multiple Tables from different files with Bulk Insert in MS SQL Server



    Earlier Features to Handle Big Data in Relational Database Server

    Advanced Data Types: BLOB (Binary Large Object) or Text/CLOB (Character Large Object) in MySQL or MS SQL Server

    How to Create and Insert to a Table with a Column of Large Text Data or Image Data in a Relational Database Server:

    Text Data Type in MySQL
    What is TEXT data type in MySQL
    Difference between blob and clob datatypes
    Blob Data type in MySQL
    how to insert blob and clob from files in mysql
    how to insert text column in mysql



  • Trouble Shooting Error Resolutions When Set Up SQL Server with ODBC/JDBC

  • For Those Who Want to Use CLR Table Function to Create a Table in SQL Server -- This is NOT For CIS492/593
    How to Set Up ASP.NET with SQL Server
    How to Debug CLR UDF, CLR UDT, CLR TVF













    Lab 2 on JSON Data Processing

  • Lab2 on JSON Transformation
  • FAQs on Lab2 on JSON Transformation
  • Transformation of JSON to a Relational Scheme in the First Normal Form and the Third Normal Form (From the CIS530 Database System Lecture)

  • Notes and Corrections:

    Note that you need to transform business.json file only for Lab2, not all of 5 json files from the Yelp site.

    Yelp Business Data Set for Lab2:

  • JSONDATA zip file

  • Yelp data set and documentation for JSON File Structures

    You can directly download the most recent data sets from Yelp site:
  • Yelp Site for Full Data Sets for Big Data Project
  • Yelp site documentation for JSON File Structures
  • Yelp site for Data Description

    Yelp Full Data Set Also Avaliable in the Big data Lab below:

    Zip file for 5 JSON files from Yelp data Challenge 2017 (Or Download directly from the Yelp site below for 2020 data sets!)

    Note:
    You Have to Use 7-Zip to unzip the zip file. Some other compression software might not be able to unzip correctly.

    The full JSON data files from the Yelp site might have a few incorrect JSON syntax detected in the data file or invalid line feed, which is common. Correct them before processing if detected.
    See FAQs for how to correct invalid JSON data

    Invalid JSON format handling:
    If there is any invalid data format is found in the input file, you can change it to the correct JSON format. For example, $$ in the “Price Range” key value pair in your input json file, the value $$ is not in quotes and this will cause to fail.

    Suggested Solutions:
    Replace $$ with 2 (meaning the price level is 2 in scale 1 - 5) in the file and try to parse the corrected file in your program.
    If , is missing between objects, add it.

    The given OneBusiness.json file and business100.json file in JSONData.zip are the corrected files.








    Lab 3 on Twitter Logging Data to NoSQL Database MongoDB (Extra Credit for This Semester)


  • Lab 3

    Twitter Logging Structure in JSON

    Twitter API 1.1 is depreciated and wont be available for new developers:
    Twitter API 1.1 is depreciated

    The New Structure for the 2.0 API for tweets: the New Structure for the 2.0 API for tweets
    FAQs to Get Twitter Developer's Account

    Note That the Recent Changes of the Twitter Site Seem to Affect Their Response Time to Process Applications for Developer's Account. 
    Due to the Delay on the Twitter Site Response Time, You Need to Apply ASAP to Collect the Twitter Streaming Data in Time


    Twitter Stream Data Collection

    Collect at least 10,000 Tweets Talking about either One of the Following Topics of Your Choice. Add More Related Keywords As Needed to Your Chosen Topic to Collect As Many Related Tweets Possible:

    Suggested Topics:

    1. Any New Major Movie or Product That Was Released Recently if Any (For example, IPhone - IPhone 12, IPhone Mini)
    Or
    2. Any Major News (For example, Russian Invasion to Ukraine)
    Or
    3. President or Any Person of Interest, or Any Two Candidates in an Election
    Or
    4. Covid, Corona Virus, Covid-19
    Or 5. Any Topics of Your Interest as long as there are big enough to collect more than 20,000 Tweets


    Twitter Data Collection Setting Up: (The Twitter Data Collected will be used for Lab3 and Can Be Used For Your Final Project later)


    NOTE that to collect the Twitter stream data in real time, you need to apply for their developer’s account in the Twitter Deveoper's site
    and get a permission to get credentials for a token and keys. This process usually takes 3-4 days or one week for Twitter to respond.

    You HAVE TO start your application ASAP. Don’t wait until the last day.


    See the example project and sample codes for the step by step procedures for this. Read everything posted in the links below.


    Apply a Twitter Developer's Account ASAP for Twitter Stream Data Collection ! Start ASAP Since it Will Take a Week to Obtain a Twitter Developer's Account !!

    For Your Twitter Developer's Account Application, Choose the most Common Account type. Do not choose an Academic Reserach Account (You Will Be Asked More Questions).

    Do NOT Blindly Copy Those Sample Answers in the Class Webpage for Your Answers ! Rephrase/Modify Them in Your Words For Your Case.

    See Examples of the Answers for the Questions from Twitter to Get a Twitter Developer's Account

    (This is For an Academic Research Account, which You Don't Need to Apply) See Sample Answers for the New Questions from Twitter to Apply Twitter Developer's Account

    FAQs to Get Twitter Developer's Account


    For Your Twitter Account Application
    For the Project Site if Asked, Provide Your Lab3 Specification above.
    you Can also Provide the CIS612 Project Site and Research Project Description for Big Data and Data Scientist

    CIS 612 Project Site to Provide in Your Application for Twitter Developer's Account
    To Answer with Sample Research Project Description for Big Data and Data Scientist



    Twitter API 1.1 is depreciated and wont be available for new developers:
    Twitter API 1.1 is depreciated

    The New Structure for the 2.0 API for tweets: the New Structure for the 2.0 API for tweets


    How to Collect a Twitter Stream to a JSON File then Insert to MongoDB

    Twitter Logging Structure in JSON
    Note some Tweets Don't have the Retweet part of info. Try to Collect Retweets Together.


    How to Get Twitter Stream in JSON file:

    Step by Step Guide on How to Get Twitter Stream in JSON file in Python (After the Tweepy API Version 4.0 as of Spring 2022) NEW POST !! by TA Yixi Luo
    ****************
    Step by Step Guide on How to Get Twitter Stream in JSON file in Python (Before The Tweepy New Version 4.0)


    How to Get Twitter Stream in JSON file then Insert into MongoDB:

    Tutorial on How to Get Twitter Stream to Insert into MongoDB in Python As of 2021 Before the New Tweepy Version 4.0

    Note: This tutorial uses last year (2021)'s Tweepy Streaming Library. It has been upgraded to a new version as of early 2022. See The Tutorial for the New Version of Tweepy Streaming Lib Above
    Blindly copy and paste of the codes in this tutorial wouldn't work because of the new version of Tweepy Streaming Lib




    Other Related Documentations and Examples from Twitter sites and Python for Twitter

  • Twitter site Documentation for Developer
  • How to Get Twitter Stream in Python
  • How to Get a Twitter Stream in Python
  • code examples for collecting Tweets



    If You Failed to Get a Twitter Developer's Account, Use this Twitter Raw data set in this site for This Lab
    Kaggle: Clean Raw JSON Tweets Data site
    RawJsonTwitterData.zip










    The Recent Changes of the Twitter Site Seem to Affect Their Response Time to Process Applications for Developer's Account. 
    Due to the Delay on the Twitter Site Response Time, Lab3 on Twitter Logging data Has Been Replaced by Lab3_1 on Yelp Data Processing.

    Lab3_1 with Yelp Data Set:

    Semi-Structured Data Processing with Semi-Structured Database Server MongoDB:

    1. Creating a Semi-Structured Database - MongoDB Collections in a MongoDB Server for business.json and review.json files from the Yelp site
    2. Writing Aggregation Pipelinig

    Lab3_1 on Writing Aggregation Pipelinig on Mongo DB Collections from json data files from Yelp site

    Make Sure to Use All the Documents in business.json to create a Collection named business for this Lab

    Data Sets for Lab3_1:
    Zip file for 5 JSON files from Yelp data Challenge 2017 (They may not be in valid JSON format) or Download directly from the Yelp site below for 2020 data sets!
    Note:
    If You are having a corrupted zip file error for 2017 Yelp Data Set, use 7-Zip to unzip.
    You can directly download the most recent data sets from Yelp site:
    Yelp Site for Data Sets


    Zip file for One business data and 100 busineess from Lab2 (in valid JSON format)



    If you have invalid JSON file problem, the JSON file might need to be corrected through the JSON validator to detect the errors first then correct the JSON syntax to be imported. The data files in the Zip file are all coming from the yelp site. They sometime have the incorrect JSON syntax, which need to be corrected to make it work. 
    FAQs On Lab3_1 on Yelp Data with MongoDB
    See the example of how to correct an invalid json file below. Scroll down for the part. FAQs for JSON Processing in General from Lab2 FAQs

    More FAQs on Lab3 on MongoDB:
    Q:
    Can we use the 2019 dataset (JSON) from yelp website for our Lab 3_1? Or we should only use 2017 JSON data?
    And Do we need to create CSV for each query as well for the count result?

    A:
    Either data set is ok. Some people have a Zip error with the 2017 data set. If then use the 2019 data set from the Yelp site. Conversion to CSV file is not required for This Lab3_1.






    Lab3 Mongo DB Guides:

    Mongo DB Setup:

    MongoDB Installation:
    MongoDB Download and Installation
    MongoDB Installation Guides

    Full Menu for MongoDB Download
    MongoDB Enterprise Download
    MongoDB GUI Client Compass Download


    Full Menu of MongoDB Manuals for Getting Started
    MongoDB Cilent: Mongo Shell (mongosh) to start
    MongoDB Shell to Connect
    MongoDB Cilent: Writing Scripts for Mongo Shell
    How to Install pymongo driver and Connect to MongoDB Server in Python application as a client

    See MongoDB Lecture Notes Section for MongoDB CRUD Queries and more details

    How to Import Data file to MongoDB
    Mongo import


    Sample Runs of Mongo DB Queries

  • Sample Project Codes: Processing JSON Data in MongoDB to SQL Table in Python

  • How to Import a json file to MongoDB

    How to Use MongoDB in Python for CRUDE: DB/Collection Creation, Insert, and Querying

    How to Import a JSON file to MongoDB in Python

    Example of MongoDB Aggregation Pipeling for Word Count

    How to Save MongoDB Query Results into a variable

    How to Save MongoDB Query Results











    Lab 4: Text Analytics (Text Mining) with Information Retrieval and Natural Language Processing Methods


    Lab 4 on Document Vectorization with TF-IDF for Document/Webpage Clustering for Categorization
    -- (For Those Who Have Already Taken CIS660, Choose Lab4_3 as Lab4 !)

  • Simple Lab 4 on Webpage Similarity Measure on (User Given) Topics (Simpler Version)

  • Lab 4_2 on Cosine Simiarity Measure for Webpage/Document Clustering for Categorization by Topics

  • Note that Building Inverted Index Is NOT Required for Lab4. So Document Vectorization Based on TF-IDF Can Be Done with the Frequncy Count Only with Stopword Removal without Document Fequency
  • FAQs for Lab 4



  • Important Changes for Lab4:

  • Required for Lab4 for CIS492 abd CIS593:
  • Note that Building Inverted Index Is NOT Required for Lab4.
    So Document Vectorization Based on TF-IDF Is Required with the Frequncy Count (without Document Fequency) Only with Stopword Removal

  • Extra Credits for Lab4 for CIS492 and CIS593:

  • 1. Building an Simplified Inverted Index in a SQL Server is an Extra Credit (50%) for Lab4

    For Inverted Index to Build, You Can Simplify to One Table with (Term, Doc#, TermFreq)

    2. Building a Full Inverted Index Either in a SQL Server or MongoDB is an Extra Credit (100%) as in the Lecture Note as below:

    Dictionary table (Term, TotalDocsFreq, TotalCollectionFreq) and Posting Table(Term, Doc#, Term_Freq)

    3. Build Inverted Index with NLP Pipelining (150%) for each Sentence to Extract Context Aware Information with POS or/and NER Tagger and Store them either in SQL Server or MongoDB for Retrieval Later




    You Can Choose to Extend One of the Labs Below with a Big Collection of Documents as a Final Group Project !

  • Lab 4_2 on Cosine Simiarity Measure for Webpage/Document Clustering for Categorization by Topics


  • Extra Credit Lab 4_3 on Information Retrieval Methods for Content Based Document Search Engine


  • Lab 4-3 on Content Based Document Search Engine with a Collection of Union Address


  • Example of Inverted Index on State Union Addresses (Note that the Structures of the Index Tables are a little Different in the Example)



    Text Preprocessing Library in Python SpaCy:

  • Python Text Processing Libs for Text Analysis
  • Liquistic Modules in Python SpaCy
    Lemmatizer in Python SpaCy
    Liquistic Modules for Tokenization, Stemming, Lemmatization in Python SpaCy
    How to Code Liquistic Modules like Lemmatizer in Python SpaCy
    Python Example for Basic Text Processing






    Lab 5 on Classification with Machine Learning:


  • Lab 5 on Classification with Machine Learning

  • Requirements for Lab5:

    - Repeat Classification with at least two Different MLs
    - For Each ML, find the best input parameters to generate the best performing model
    For Example, with SVM, try with different kernel functions to find the best fit(model) for the given data

    Optional for extra credit (10%)
    - K-folder Cross Validation with K = 5 or 10


  • FAQs on Lab 5 on Classification with Machine Learning

  • For Lab 5, You Can Choose to Use Any of the Following Data Sets:

    Note that the Data files and Data Description Files are text files, you can open them as a text file with any text editors like Notepad or Notepad++

    1. Adult Profile Data Set to Predict Income Level < 50k or NOT from Adult Data set

    UCI Site for Census Adult Profile Data Set and Data Description for Binary Classification
  • Census Adult Profile Data Sets of Lab 5 on Classification


  • 2. Wine Data Set with either DT or NN to Predict Wine Quality in Scale 1 - 10

    UCI Site for Wine Data Set and Data Description for Multi-Class Classification
  • Wine Quality Prediction Data Sets of Lab 5 on Classification


  • 3. Patient Data Set for Breast Cancer Prediction for Binary Classification to Predict whether the Patient has a Breast Cancer or not
  • Health Data Set: Patient Data for Breast Cancer Prediction for Binary Classification



  • ML Algorithms -- ANN or SVM Reqire Data Preprocessing:
    Normalization for Any Numerical Attributes
    Binarization (One Hot Encoding) for Any Categorical Attributes

       Tutorial for Data Preprocessing Normalization and One Hot Encoding for ANN or SVM

    Categorical Data transformation Methods with Binarization (One Hot Encoding)

    Data Preprocessing Methods for ANN or SVM



    You Can Choose Your Own Data Set to Extend Lab5 for Final Group Project !

    Data Set Repository for Classification:

    Kaggle Data Set Repository
    UCI Data Set Repository for Classification
    Health Data Set Repository for Classification
    Socrata Open Data API





    Extra Credt Lab 5_1 on Sentiment Analysis with Machine Learning for Classification (Not Required for Everyone):
    Extra Lab 5_1 on Sentiment Analysis with Machine Learning for Classification:

    Data Set: Choose a Review Text Data Set Obtained from Social Network sites: Twitter or Yelp Review Data Set

    Classification Goal:

    1. For each review text obtained from Twitter texts, Derive a Preditive Model to Predict (Classify) Whether it is Positive or Negative Review.
    2. For each review text in Yelp Review Data Set, Derive a Preditive Model to Predict (Classify) the scale of the Review in 1 - 5. (star)




    Useful Big Data Analytic Tools

    Choose your System/Tool/Platform to Set Up and Get Used to:

    Python Analytics Tools and Tutorials :

    Machine Learning :

  • Python Scikit Learn
  • Python Scikit Learn for Data Preprocessing
  • Python Numpy Tutorial

  • Other Machine Learning Platforms:

  • Pandas: The Python Machine Learning library
  • Anaconda Machine Learning Platform for Python, R
  • Pandas Python Tutorial Codes
  • Keras: The Python Deep Learning library for Tensorflow, CNTK

  • Basic R Tutorials :
  • R Studio basic Tutorial
  • Examples to start with R Stat Tool
  • Examples of Basic R Stat Tool with Helpful References







  • Extra Credit Lab 6 on MapReduce and Hadoop (HDFS): Extra Credit Lab (Not Required for Everyone)

  • Lab 6 on MapReduce and HDFS


  • Hadoop Set Up Instruction Sites:
    Hadoop single node setup
    Hadoop Cluster setup
    Map Reduce Tutorial on Hadoop


    Lab Guides: Newest on the Top
    Hadoop Installation: How to Fix When Data Nodes are not Running (2018)
    Help Site on How to Fix When Data Node are not Running (2018)
    Procedure to How to Execute MapReduce in Eclipse to Run a Wordcount Job (2018)
    Procedure to Install Hadoop and Run a Wordcount Job on Window 2010 (2017)
    Procedure to Install Hadoop and Run a Wordcount Job on Window 2010 - Permission issue (2017)
    Procedure to Install Hadoop and Run a Wordcount Job on Mac (2017)
    Guideline4 for Passwordless SSH for Lab4_1 (2016)
    Guideline2 for Lab4_1 (2015)
    Guideline1 for Lab4_1 (2014)

    A good Instruction Site for Installing and Running Hadoop
    Installing Hadoop

    If you have a trouble installing Hadoop from the above site with not seeing the name node, You need to delete the temp files created in standalone mode and reformat the namenode.
    For setting up a passwordless SSH, see Ganesh's Lab4 below as well.
    Installing Hadoop and running a wordcount job by Ganesh VAVILAPALLI
    Video for Installing Hadoop shared by Prashant Patel
    Trouble Shooting Tips for Installing Haddop on VM
    Guideline3 for Lab4_1 on Mac (2013)


    For Those who Want to Set Up Your HDFS Cluster On EC2 Amazon Cloud, See the Cloud Section at the end of the Class Lecture Note Section for a Student Account.




    Supporting Contents for Labs :

    More to Come !
    For Installation Guides:
    SQL Server Installation Guide site
    Read the Installation Guides FIRST before starting downloading ! See the guide site for More details !

    How to Create a Web Application with Java Based Application Server with MS SQL Server:
  • Set Up Instructions for Java Based Application Server with MS SQL Server
  • How to Create Java Servlets for CRUD Operations
  • How to Create Java Servlets for CRUD Operations


  • Amazon Cloud:
    Amazon RDS (Relational Database Service)
    Amazon Elastic Cloud Computing (EC2) for Web service
    How to Create Amazon Virtual Hosting (EC2) with a Web Server and Amazon Database Server (RDS)

    Microsoft Cloud Azure:
    Trial account for Microsoft AZURE Cloud
    How to Create/Retrieve a Table in Microsoft AZURE Cloud
    Sample Project on Microsoft AZURE Cloud
    How to Create/Retrieve BLOB data in Microsoft AZURE Cloud
    AZURE Cloud
    How to Create a SQL Database Server in Microsoft AZURE Cloud
    Tutorial for MS AZURE Cloud


    Useful Resource Sites:
  • XHTML on WWW.W3.Org
  • XML Version 2.0 on WWW.W3.org
  • DOM XML Parser
  • Microsoft XML DOM Parser Beginner's Guide
  • SAX XML Parser
  • Example Codes of XML Processing with XML Parser (DOM Parser and SAX Parser)
  • XML Editing with OXIGEN
  • Useful XML Resource Sites
  • Beginner's Guide to XML DOM


  • Class Lecture Notes with Tentative Schedule

    Class Chapter / Topic / Specific Objectives / Activities
    1


    Introduction to Big Data, Big Data Processing, Big Data Processing Systems, and Big Data Anaytics for AI

    Lecture Notes_1: Introduction to Big Data, Big Data Processing and Big Data Analytics

    Lecture Notes_1_2: Core Builing Phases of AI with Big Data, Big Data Processing and Big Data Analytics

    5-Min Overview on What is Big Data Analytics

    Introduction of Data Science with Big Data and Research on Big Data Analytics at CSU




    Getting to Start with Examples of Intelligent Systems (AI) with Big Data, Big Data Processing and Big Data Analytics:

    Examples of Intelligent Web Applications


    Examples of AI Applications with Big Data

    LectureNotes_2_1: Overview of Question Answering System

    LectureNotes_2_2: Overview of IBM Watson: Question Answering System

    Lecture Notes_2_3: Overview of Sentiment Analysis System


    Examples of Big Data Projects from Best CS Projects:
    Project Example_1: Sharesci - Architecture of an Intelligent Document Search Engine: sharecsi
    Project Example_2: Overview of Intelligent Document Search Engine Using Machine Learning with Natural Language Processing

    Examples of Big Data: 7.2 Million Wiki Webpage Dump in XML format to download to process

    wiki site of sharecsi for Documentation
    github site of an Intelligent Document Search Engine: sharecsi for Source Codes


    What are the Most Imprtant Topics to Learn in Big Data ? All of Them Are Covered in CIS593, CIS611, CIS612, and CIS660
    For Comparison: Topics in Big Data Specialization

    Realated Subjects to Be Covered in this Big Data Course
    For Comparison:Topics in Big Data Engineering Specialization



    2-5





    Semi-Structured Data Processing for Big Data Applications (Serverside Processing)



    Semi-Structured Data Processing Techniques in Web Applications

    Overview of Architecture of a Web Application Platform

    Lecture Note on Semi-Structured Data Model with XML and JSON


    Universal Data Exchange Formats on Internet Between Web Applications As Client and Server -- Platform Independent

  • JSON (JavaScript Object Notation) in Semi-Structured Format
  • XML (eXtensible Markup Language) in Semi-Structured Format
  • HTML/XHTML (Hyper Text Markup Language) in Semi-Structured Format
  • CSV (Comma Separate Value)/TSV(Tab Separate Value) in Structured Format from Relational (Table) Model


  • Good Example of Web Service Site
    Bad Example of Web Service Site
    Bad Example of Web Application Site




    Common Big Data Formats:

  • Millions of Webpages -- Text Documents with Markup Tags
  • Social Media Server Logging Data -- Twitter or Facebook Server Longging Data -- Millions of Semi-Structured Data in JSON or XML Documents
  • Web Server Logging Data -- Structured or Unstructured Text Documents

    Big Data Examples in Semi-Structured Data in JSON Format from Real Life Applications:

    Data sets at Yelp Site
    Yelp Data Format: Business.json
    Social Media Twitter Server Generated Data in JSON
    Structure of Twitter Server Logging data in JSON
    7.2 Million Wiki Page Dump either in HTML or XML




    Common Big data Format in Semi-Structured Data Model:

    Big data in Semi-Structured Data Models for Data Exchange Formats between Web Applications in Internet

    Web Application Architecture


    HTML/XHTML/XML Documents (a Collection of Millions of Webpages) as Big Data in Semi-Structured Data Model:

    REVIEW on HTML and XML:

    Markup Languages: HTML, XHTML, XML as Semi-Structured Data Model

    Big Data as Semi-Structured Data on Web Application Platforms on WWW


    World Wide Web Consortium (W3C) Reference Sites for MarkUp Languages:

    World Wide Web Consortium (W3C) Standard for Markup Languages: HTML/HTML5/XHTML
    HTML/HTML5
    HTML/HTML5 Syntax by W3 Consortium
    XHTML1 XHTML2 is on the way !


    1. HTML/XHTML:

    REVIEWS on Markup Languages and Basic Web Technologies:

    See See CIS408 Class Lecture Notes for Tutorials on HTML, JavaScript

    HTML
    HTML Basic Summary
    Difference between HTML/HTML5 and XHTML

    HTML References:
    HTML Element List
    HTML Attribute LIST
    Global HTML Attribute LIST







    Examples of Big Data in Webpages as Semi-Structured Data Model:

    Webpage Processing Techniques (for a Large Collection of HTML source files or in XML data files) with Semi-Structured Data Model:

    Examples of Big data in Web data (either in a large collection of html source files or in XML data files) to process to build an Intelligent Queston Answering System:

    Bad Example of Web Service Site to Collect Data from
    Information Service Site for the US President State Union Addresses in HTML
    Trip Advisor Site with Hotel Reviews in HTML
    Example of a Wiki Webpage
    an Example of Wiki Webpages in a HTML document
    7.2 Million Wiki Page Dump either in HTML or XML


    Big Data Processing Technologies for a Large Collection of HTML Pages as Semi-Structured Model


    Lectures:

    Semi-Structured Data Model:
    Class Note_10: Lecture Notes On Semi-structured Data Model and XML

    HTML/XHTML/XML in Semi-Structured Data Model - Document Object Model (DOM):


    Running Example of Webpage Processing Method to Extract Information from Large Corpus of Webpage Documents:

    The websites for the collection of the US President State Union Addresses
    Example of an HTML Document as Source Text File to Process



    Document Object Model (DOM) for Semi Structured Data Processing

    DOM is the Implementation of Semi-Structure Data Model for HTML/XHTML/XML


    Big Data Processing Techniques for a Large Collection of HTML Documents or XML Documents

    HTML/XHTML/XML in Document Object Model (DOM)

    DOM Lectures:

  • HTML in Document Object Model (DOM):

  • Class Note_3: Intro to Document Object Model (DOM) for HTML and JavaScript
    Class Note_3_1: Example Picture of DOM Tree of Web Browser for HTML and JavaScript
    Example of Document Object Model (DOM) for HTML and JavaScript

    Node Types of DOM Tree: (From W3C Schools Tutorial on HTML DOM)

    DOM Document Node Type for Document Element HTML
    DOM Element Node Types and Methods/Properties
    DOM Attribute Node Type and Properties
    DOM Text Node Type

    HTML DOM Reference List in JavaScript:
    HTML DOM Document and Method List
    HTML DOM Element and Method List
    HTML DOM Attribute and Method LIST
    Important Note:
    DOM Query Result in Collection vs List


    MDN: HTML DOM Intro
    MDN: DOM Parser for HTML/XML
    MDN: Useful Example Sites of DOM Objects
    MDN: HTML DOM API


    Official DOM References on W3C:
    Official Document Object Model (DOM) Specification for HTML/XHTML/XML by W3 Consortium


    IMPORTNT NOTES !!!

    Note that the code examples in JavaScript in this section Are Only for Client-side Processing for Dynamic Webpage Changes Where a Web Browser Automatically Parses and Builds a DOM tree for the Html file to Execute.
    This client side Javascript examples are only for your understanding.
    All the labs and Final Project in this course require Server-side Data Processing in any modern Server-side (script) languages such as Python, C++/C#, PHP, Java, or Node JS (Server-side JavaScript with File I/O)

    Code Examples of DOM Tree :

    Example of DOM property InnerHtml
    Example of How to Find DOM Objects: getElementByTagName
    Example of How to Access DOM objects and values: childen element count
    Example of HTML DOM Object with button
    Element.className

    Example of How to Access DOM Atrributes: namednodemap: getnameditem
    Examples of HTML DOM Object creation
    Examples of HTML DOM Object Create and Append
    Example of How to Create Text Node
    How to create DOM Objects and Replace: replaceChild



    IMPORTNT NOTES !!!

    All the labs and Final Project in this class require Server-side big data processing in any modern Serverside (script) languages such as Python, C++/C#, PHP, Java, or Node JS (Server-side JavaScript)

    When You Build a Server-side Application with any Programming Language to Process HTML/XML Documents with DOM and XPATH for an Big Data Application,
    You Have to Call a DOM Parser to Build a DOM Tree for Each HTML/XML Document !

    In any other seperate script/programming lanaguages like Python, Java, or Node JS, you have to use(call) a DOM Parser that will build a DOM Tree from your input document like html or XML file.



    Useful DOM Tree Query Features for Big Data Proceesing:

    XPATH: See the XML Section Below for the XPath Lectures !
    Setting Up lxml in Python for DOM with XPath and Code Examples ***************
    XPath in JavaScript

    Query Selector with CSS, JS RegExp:
    Document.querySelector with CSS Selector or Rex
    get elements by selector: querySelector with CSS Selector in Python
    Review on CSS and JS RehExp:
    HTML CSS Rules to Select Elements
    Letcture Notes JavaScript Basics (See Page 18 - 20 on Regular Expressions)
    JS RegExp
    RegExp: ignoreCase property of i modifier


    DOM with XPATH Set Up in Any Script/Programming Languages for Server Side Applications

    Note that All the RECENT Web Browsers Have DOM and XPATH Features in their Debugger by Default.
    You don't Need to Do Any Extra Set Up to Use XPATH in your JavaScript in a Client Code for HTML Processing in your Webbrowser.

    Any Modern Script/Programming Languages Have the DOM Parser and XPath Features. You Need to Set up to Use !
    For the Labs to Write a Server Application in this course, You Need To Add DOM Parser Related APIs for Your Scripts/Programming Languages to Do:

    1. Call a DOM Parser to Build a DOM Tree
    and
    2. Use XPATH Query Methods to Retrieve DOM Objects to Extract Information (Tag Names or Text) in the Application Codes as in the Examples below.


    Useful Webpage Processing Tools with Built-In DOM Parser and XPath Query Feature:

    Setting Up lxml in Python with DOM Parser and XPath
    Beautiful Soup Documentation for Document Processing for HTML/XHTML, XML in Python
    DOM Example: Webscrapping with Python BeautifulSoup

    Selenium-Python: Setting Up DOM with XPath for Python
    Selenium-Python: XPath API to Locate Elements in Python


    Code Examples of DOM and XPath for Document Processing in Python:

    Python XPath Code Examples
    Webscapping with DOM, XPath for Python with Beautiful Soup (From the Stanford Class)








    Think about What Would Be a Better Way to Send the Contents of the Webpages of the Collection of State Union Addresses to Application Servers ???


    XML (eXtensible Markup Language) Documents in Semi-Structured Data Model

  • XML in Document Object Model (DOM)

  • Lectures:

    Semi-Structured Data Model for Universal Data Exchange Formats

    Lecture Note on Semi-Structured Data Model with XML and JSON

    Better Way to Send the Contents of the Webpages of the Collection of State Union Addresses to Application Servers !


  • XML (eXtensible Markup Language):
  • XML as Universal Data Exchange Format in Semi-Structured Model

    XML is an Universal Data Encoding Protocol to Seperate Data/Database (Data and their Scheme) from Presentation (Displaying instruction Codes) in Html

    Class Note_5: Introduction to XML

    Well-Formed XML Syntax

    Tutorial for XML with Examples of XML Data
    simple.xml file with XSLT file: Simple Example of XML Data file with XSLT Instruction file
    simple.xsl file: Simple Example of XSLT Instruction code file for XML Data

    Tutorial: How to Use XML
    Tutorial: XSLT (XML Stylesheet Language) Transformation
    Example of Data in XML and Presentation Instructions in XSLT
    XML DOM Tree Example
    Data Processing for XML DOM




    1. Information Extraction Techniques from a Large Collection of XML Documents as Semi-Structured Data in DOM

    On Recieving XML Data in a Text Format from a Server site, Client Site Has to Convert a XML Data to Objects in DOM Tree Using XML Parser !!
    XML Parser in JavaScript to Convert a XML Data to a DOM When a Client Recieved a XML Data in a Text Format in Javascript for Client


    Data Processing for XML DOM
    Example of XML Document Object Model (DOM) Tree for XML Data



    Basic XML Data Processing Examples in JavaScript with AJAX in HTTP in a Client Site:
    XML File Example: cd_catalog.xml
    AJAX XML File Processing Examples
    AJAX Examples - XML File Processing
    AJAX XML Processing Example




    Information Extraction Methods for Big Data Applications

    How to Select(Query) XML DOM Elements

    Useful DOM Tree Query Features:

    XPATH:

    Lecture Notes On XPATH

    Tutorial:
    XPATH To Query (Serach) XML DOM Tree
    XPATH Syntax
    Tutorial: Summary on XPath with Examples
    W3 Tutorial: XPath Quick Syntax Summary


    Important Notes on Information Extraction from HTML or XML Documents for Labs:

    When You Build an Application Server to Process HTML/XML Documents in DOM with XPATH in a Server-side (Script) Language, You Have to Call a DOM Parser to Build a DOM Tree for Each HTML/XML Document !

    In any other seperate script/programming lanaguages like Python, Java, or Node JS, you have to use(call) a DOM Parser that will build a DOM Tree from your input document like html or XML file.

    Note that You have to Use a DOM Parser to Build a DOM Tree from your Input HTML or XML file to be able to use XPATH query !!
    In any script/programming lanaguages like Python, Java, you have to use(call) a DOM Parser that will build a DOM Tree from your input document like html or XML file.


    DOM Parser with All the DOM Methods and Query Features/APIs like XPATH is available in any modern programming languages - Python, Java, C++/C# for Big Data Applications


    XPath API is available for Querying HTML DOM Tree as well as XML DOM Tree either in Web Browser Debugging Tool or any Programming Languages - Python, Javascript, Java, C#, PHP:

    Example Code for Setting Up DOM with XPath in Python
    Tutorial: XPath Syntax


    XPATH Set Up in Any Script/Programming Languages for Server Side Applications

    Note that All the RECENT Web Browsers Have XPATH Features in their Debugger by Default.
    You don't Need to Do Any Extra Set Up to Use XPATH in your JavaScript in a Client Code for HTML Processing in your Webbrowser.

    The Labs in This Course is to Write Tasks to Be Done in an Application Server, You Need To Add DOM Parser Related APIs for Your Scripts/Programming Languages
    Any Modern Script/Programming Languages Have the DOM Parser and XPath Features. You Need to Set up to Use !

    To Do 1. Call a DOM Parser to Build a DOM Tree for Each Document in Your Data Collection
    and
    2. Use XPATH Query Methods to Retrieve DOM Objects to Extract Information (Tag Names or Text) in the Application Codes as in the Examples below.





    XML DOM Processng Examples in Different Server side Programming Languages:

    Example Codes of XML Processing with XML Parser in Java

    Code Examples of Conversion from XML to Table in PHP

    MDN Site for Code examples of DOM XPATH in JavaScript
    Examples of Mobile APP On How DOM Parser and SAX Parser Work to Process XML data (Examples in Mobile App)


    Any Modern Script/Programming Languages Have the DOM Parser and XPath Features. You Need to Set up to Use !

    XML API for Server side Applications: DOM Parsers to Download for Programming in JAVA or Other Languages:

    SAX XML Parser
    DOM XML Parser
    Which DOM Parser jar file Should I Download
    MDN: DOM Parser for HTML/XML


    Useful DOM Parser and XPATH APIs with Example Codes in Python:

    Setting Up DOM with XPath with lxml in Python
    BeautifulSoup Documentation for HTML, XML in Python
    Python XPath Code Examples
    Webscapping with DOM, XPath for Python with BeautifulSoup (From the Stanford Class)

    Selenium-Python: Setting Up DOM with XPath in Selenium-Python
    Selenium-Python: XPath API to Locate Elements


    Example Codes of DOM and XPath in JavaScript:
    Note that Although the code examples here are in JavaScript, but All the DOM API and XPATH Query Features are available in any modern script languages - Python, Java, C++/C#

    MDN Site for DOM with XPath in JavaScript
    DOM with XPath Setup Guide in JavaScript with Node JS
    List of XPath Methods in JavaScript with Node JS


    MDN: DOM Parser
    MDN Site for Code examples of DOM in JavaScript
    MDN: How to Use XPath in JavaScript









    2. Big Data Processing for a Large Collection of Webpages as Unstructured Text and DOM concept - Parsing and Searching:

    Webpage Processing Tools/APIs/Features:

    Python BeautifulSoup Documentation
    Webscrapping with Python BeautifulSoup

    Regular Expression Rules in JavaScript (See Regular Expressions part at the end)
    JS RegExp Summary (Any Other Luanguages Also have Similar RegExp Features)
    RegExp: ignoreCase property of i modifier




    3. Features of Relational Database Server for Big Data Processing

    Automatic Table Creation in a Database Server:

    pyodbc(Open Database Connectivity) in Python to Connect MS SQL Server :

    Installation pyodbc to Connect to MS SQL Server in Python
    pyodbc to Connect to MS SQL Server in Python
    How to Connect to MS SQL Server in Python
    pyodbc: Examples of ODBC in Python

    How to Create a Table in a SQL Server from CSV/TSV file

    How to Create a Table with Bulk Insert with MS SQL Server
    Example of a Script to Create Multiple Tables with Bulk Insert in MS SQL Server
    Example of how to import csv file to mysql table
    Example of how to export table to csv in mysql



    Tutorial Site: How to Extract Data
    Tutorial Site: Text to Speech


    Earlier Features to Handle Big Data in Relational Database Server

    Data Types in SQL Server: BLOB (Binary Large Object) or Text/CLOB (Character Large Object) in MySQL or MS SQL Server

    How to Create and Insert to a Table with a Column of Large Text Data or Image Data in a Relational Database Server:

    Difference between blob and clob datatypes
    What is TEXT data type in MySQL
    Text Data Type in MySQL
    how to insert blob from a file in mysql using-loadfile-method
    insert-retrieve-file-image-as-a-blob-in-mysql



    Industry Example of Database (Semi Structured in Table format) of Extracted Information :

    Google Big Table for Big data processing












    Industry Example of XML data files as Big data (a big XML data file to encode a database) to process to build an Intelligent Queston Answering System:
    7.2 Million Wiki Webpage Dump in XML format to download for a Big data Project

    W3 Consortium on XML with DOM and XPath References:
    Complete List of DOM with XPath in www.w3.org













    4. JSON (JavaScript Object Notation) as Semi-Structured Model:

    JSON (JavaScript Object Notation): http://www.json.org/

    Big Data in JSON Format in Real-life Industry Applications:

    Data sets at Yelp Site
    Yelp Data Format: Business.json
    Structure of Twitter Server Logging data in JSON


    Big Data in Semi-Structured Data Model in JSON

    JSON Lectures:

    Class Note_10: Lecture Notes On Semi-structured Data Model and Comparisons to HTML, XML, JSON, Relational Table

    Class Note_11: Quick Tutorial on How To Process JSON Data


    JSON Syntax:
    What is JSON ? How to Exchange and Process JSON Data
    JSON Syntax vs Javascript Object Syntax for values

    Example of json data format of a real life application : Yelp One Business Data Format


    How to Exchange JSON Data for Big Data Processing:

    1. Parse: On Recieving JSON data (.json file): Parse JSON Text data to Convert to ==> JSON Objects in Applications in any language to Extract infomation (scheme and values) to Create a Database

    2. Stringify: To Send Database infomation: Encode(Build) JSON format in JSON objects from Database then Stringify to Convert JSON Objects to ==> JSON Text data (.json file) to Send


    How to Exchange and Process JSON Data: JSON.parse and JSON.stringify
    JSON.parse
    Processing Date in JSON
    JSON Data Processing -- JSON.stringify

    JSON Object
    Accessing Properties(Names) of JSON Object
    Accessing Property Values of JSON Object
    Accesing and Changing JSON Object
    Delete Object in JSON
    JSON Array - Nested Array


    JSON Processing Examples:

    Client side in a HTML with Javascript
    JSON Data Processing with HTML in a Client Side

    Handling in Server side application codes:

    Python:

    1. To Parse JSON data to Convert to ==> Python objects: json.loads()
    2. To Stringify Python objects to Convert to ==> JSON text: json.dumps()

    Intro to JSON Data Processing in Python
    Example of JSON Data Processing in Python

    JSON Data Processing with PHP in Server Side
    Lecture Note6_5_5: On How to Process JSON data in a Mobile App






    Object Relational Mapping (ORM)

    ORM Mapping from Semi-Structured Data to Structured Realtional Database Scheme:

    Transformation of Semi-Structured Model to a Relational Model in the Third Normal Form



    Reviw:
    Realtional Model and Constraints

    The First Normal Form Rule in Relational (Database) Model

    Design Rules of Relational (Database) Model



    Official Reference to JSON
    http://www.json.org/










    5. Big Data: Unstructured Text Data Processing for Information Extraction

  • Completely Unstructured Text Documents -- Social Media Messages, Electronic Books (in pdf)

  • Introduction to Text Processing for Information Extraction for Text Analysis Applications:

  • Inverted Index: Efficient Index for Unstructured Text Processing

  • Class Note_18_1: Lecture Notes On Inverted Index in Google (Stanford)


    Example of Inverted Index Tables for Text Analysis in State of Union Addresses


    Google's Invered Index on Their Big Data - Entire World Wide Webpages:

    Google NGram Viewer : Inverted Index for all N gram words



    Basic Text Preprocessing with Natural Language Processing (NLP) APIs for Implementation of Text Analysis Applications:

    Text Cleaning for Preprocessing:
  • Common NLP Preprocessing Tasks to Be Done for Text Analysis
  • Example Project on Twitter Text Mining and Sentiment Analysis in R


  • Problems (Limitations) with Lexicon Based Text Analysis:

    Phrase (N-Gram Word) Identification
    Can NOT Identify Relationships among Terms - Synonyms (Similar Meaning) or Antonyms (Opposite Meaning) of Terms
    Can NOT Identify Different Meanings of a Same Term by the Different Context of a Senetence - Polysemy Problem !
    Can NOT Identify New Terms or Changining Relationship Between New Terms and Old Terms


    IMPORTANT NOTES !!!
    Note that You Have to Decide Which Text Cleaning Preprocessing Should Be Applied or Not Applied Depending on Type of Text Analysis

  • For TF-IDF Based/Lexicon Based Text Vectorization (to be covered in Text Analysis section below), You Should Apply All the Text Cleaning Preprocessing.
  • On the other hand, for Context Aware NLP Methods Such as POS and NER Tagging, or Word2Vec Training for Word Eembedding, You Should Not Remove Each Sentence Strcuture, So Should Not Apply These Common NLP Cleaning/Preprocessing -- Stopword Removal, Stemming/Lemmarization Which Remove Structures of Sentences.



  • Stemming/Lemmatization:

  • Common NLP Preprocessing Task: Stemming-Lemmatization in NLTK python
  • Common NLP Preprocessing Task: Stemming-Lemmatization from Stanford NLP Group



  • Lectures:

    Some Solutions for the Problem that Lexicon Based Methods Can NOT Identify Different Meanings of a Same Term by the Different Context of a Senetence

    Context Aware Natural Language Processing(NLP) Methods:

    Information Extraction Methods with Natural Language Processing(NLP) Technologies for Unstructured Texts
    Context Aware (Meaning of a Word from a Sentence Context) NLP Methods for Information Extraction:

  • POS (Part of Speech) Tagging

  • Lecture Notes on POS (Part of Speech) Tagging (Stanford)
    English Grammar
    UPenn Treebank POS Tagging List


  • NER (Named Entity) Tagging

  • Lecture Notes on NER (Named Entity Recognizer) (Stanford)



    Stanford Core NLP Parser Download for POS and NER tagging (Github repo site as well)
    Stanford Core NLP Demo: POS and NER tagging Examples
    Stanford Core NLP Run Site

    Stanford Core NLP Data Pipelining for Information Extraction



    Python NLP Text Processing APIs to Build Data Pipeling for Text Preprocessing for Information Extraction :

    Python spaCy Libs for stemming, lemmatization or NLP Preprocessing Methods below !

    Introduction to Text Processing in Python NLP Lib spaCy:

    Use Python NLP Lib called SpaCy to build a data pipelining with NLP APIs for Text Analysis

  • Download and Installing spaCy
  • Introduction to spaCy101 for Text Proceesing with NLP Methods
  • Python spaCy Objects for Text Processing in Data Pipelines
  • spaCy: Building Data Pipleines for Text Proceesing Tasks with NLP Methods
  • spaCy NLP Features
  • spaCy NLP Method: Rule Based Matching


  • Tutorials of spaCy for Text Proceesing with NLP Methods



  • Other Useful Core NLP Software to Download to Use

  • NLTK (Natural Language Tool Kit)
  • Stanford Core NLP APIs (from Stanford NLP Research Group)





  • Other Parsers to download:

    Stanford NLP Software Site

    Stanford NLP POS (part Of Speech) Tagging
    Stanford NLP Parser
    Stanford NLP POS Tagger in JAVA


    WordNet Site (by Princeton) to downlaod NLP Software:

    WordNet by Princeton
    Download WordNet for Sentiment Analysis on Aspect/Feature Opinion Analysis

    WordNet Sinset Similarites for All POS Tagging for Dictionary Contruction for Text Analysis

    5-6



    Semi-Structured NOSQL Database Systems for Big Data Processing

    MongoDB:

    What is MongoDB?: Overview of MongoDB Atlas (Cloud instance), MongoDB Community (the Simplest Version), and MongoDB Enterprise (Stand alone)

    Mongo DB Setup Related:

    MongoDB Installation:
    MongoDB Download and Installation
    MongoDB Enterprise Download
    MongoDB Installation Guides on Window
    Or
    MongoDB Community Version Download and Installation


    MongoDB Clients Download - MongoDB Shell (Command-line) or MongoDB Compass (GUI)



    MongoDB Manual - Getting Started:
    Full Menu of MongoDB Manuals for Getting Started
    MongoDB Shell (mongosh)
    Download and Install MongoDB Client: MongoDB Shell (mongosh)
    Start your client: MongoDB Shell (mongosh) to Connect to MongoDB Server
    Old MongoDB Shell (mongo) Replaced by mongosh
    How to Install pymongo driver and Connect to MongoDB Server in Python application as a client

    How to Import a json file to MongoDB either in Command line Client or in Python or Java

    How to Import Data file to MongoDB
    Mongo import

    Platform Specific Set up Guides
    Node JS with Mongo DB Setup Guide
    Node JS with PostgreSQL Setup Guide
    Web Application Using Node JS with MS SQL Server and Mongo DB (From a Master Project: Building a Recommendation System usng Amazon and Walmart Product and Review Data Sets)

    Complete Mongo DB Documentation Site


    Lectures:

    Lecture Notes on Overview of MongoDB
    Mongo DB Comparison to Sql Terms and Queries

    Mongo DB Intro with Examples:
    Introduction to Mongo DB
    Mongo DB: Databases/Collections
    Mongo DB: Document Format
    Mongo DB: CRUD Operations
    Mongo DB Getting Started
    Mongo DB: How to Create a database



    Mongo DB Query Compared to SQL:
    Mongo DB Mapping Chart to Compare to SQL

    Mongo DB Basic CRUD Operations:
    Mongo DB: CRUD Operations
    MongoDB Operations: InsertMany
    MongoDB Operations: UpdateMany
    MongoDB Operations: UpdateMany with Aggregation Pipelining
    Mongo DB Operations: ReplaceOne

    Mongo DB Query: Select Operation
    Mongo DB Select Operation on embedded documents
    Mongo DB Select Operation on arrays
    Mongo DB Select Operation on array of embedded documents


    Example Runs of Mongo DB Queries


    MongoDB Aggregation with Pipelining:

    MongoDB Aggregation Mapping Chart to SQL
    Mongo DB: Aggregations
    Mongo DB: updateMany with Aggregation pipelining to add a new name and value into the existing documents (Scroll down to see the examples)
    Mongo DB Operator $unwind to unpack an array list
    Mongo DB Join Operator: $lookup
    Mongo DB Join Operator: $lookup with $unwind for array


    Tutorials in Python:

    How to Use MongoDB in Python for CRUDE: DB/Collection Creation, Insert, and Querying

    How to Import a JSON file to MongoDB in Python

    Example of MongoDB Aggregation Pipeling for Word Count

    How to Save MongoDB Query Results

    How to Save MongoDB Query Results into a variable


    Examples:
    Example of MongoDB Aggregation Pipeling for Word Count

    Application Code Examples in Python with MongoDB and SQL Server

    Example Codes of JSON Data Processing in Python with MongoDB


    Mongo DB Resources:
    Class Note_22_3: MongoDB Resources
    Mongo DB Documentation




    Semistructured Database System - MongoDB: More Advanced (Not Required for CIS492/593)


    MongoDB Selection Query:
    Mongo DB Basic Selection Query: $find operator with Condition expressions
    Mongo DB Query (select): $find Example Explained

    MongoDB Query with Comparisons to SQLs:
    Mongo DB: Query find operator basic
    Mongo DB: CRUD Operations
    Mongo DB to SQL Basic Query Mapping Chart (Comparisons to SQL in Examples)

    MongoDB Expressions for Conditions:
    Mongo DB Condition Using $expr With Conditional Statements
    Mongo DB Logical Operators
    Mongo DB Condition with Logical Operator $and

    MongoDB Aggregate Pipelining:
    MongoDB: Intro to Aggregation Pipelining
    MongoDB Query/Aggregation Pipelining Operator Mapping Chart to SQL
    Class Note_24_1: Summary of MongoDB Query/Aggregation Ppelining Exampes with Comparison to SQL
    Complete List of MongoDB Aggregate Operators for Pipelining

    MongoDB lookup Operator for Join with Aggregation Piplelining :
    MongoDB $lookup operator for Join: Join returns null matching as well (Left Outer Join Semantics)
    MongoDB $lookup operator with $unwind for Array Valued Field as Join column
    MongoDB $lookup operator with $let for Multiple Fields for Join Columns
    MongoDB let operator

    MongoDB CRUD -- More Advanced:
    MongoDB UpdateMany Operation with Aggregation Pipelining and $set as $addfields, $unset Operators
    MongoDB Add Field with Aggregate operation Pipelining
    MongoDB Add Field with Aggregate operation Pipelining
    MongoDB Add Field with Aggregate operation Pipelining


    Examples of Big Data Project with MongoDB and Yelp Data

    Project Presenation
    Code Examples with MongoDB and Angular JS


    Examples of Big Data Applications from a Previous Research Team in Big Data Anaytics Lab

    sharesci: Intelligent Search Engine in Natural Language

    sharesci: Content Based Document Search Engine on wiki pages and research papers
    sharesci: Content Based Document Search Engine with Machine Learning

    sharesci github site for Source Codes
    sharesci wiki site in bigdata lab for documentation
    sharesci Project Report for Source Codes


    10-11



    Big Data Analytics: Analysis of Unstructured Text

    Motivation:

    Lecture Notes_1_2: Core Builing Phases of AI with Big Data, Big Data Processing and Big Data Analytics

    Applications of AI:

    LectureNotes_2_1: Overview of Question Answering System (AI that Can Answer to a Question (Query) in a Natural Language)

    LectureNotes_2_2: Overview of IBM Watson: Question Answering System

    Lecture Notes_2_3: Overview of Sentiment Analysis System



    Important Lectures for Unstructured Text data Processing:
    ***************
    Class Note_18_2: Lecture Notes On Information Retrieval for Google Search engine (Stanford) ***********

    Related Earlier Lectures:

    Class Note_18_1: Lecture Notes On Web Data Processing: Inverted Index in Google (Stanford) ***********

    Class Note_18_3: Positioning Index for Phrase (N-Gram) Query (Stanford) (Scroll Down to Phrase Query Section from Slide 41 )


    Similarity Measures:

    Jaccard Similarity with Example ***********
    Cosine Similarity with Example ***********

    Similarity/Disimilarity Measures|




    Introduction to Text Processing for Implementation:

    Basic Text Preprocessing with NLP APIs:

    Twitter Data Analytic Processing with Natural Language Processing (NLP) methods

    Text Cleaning for Preprocessing:

  • Common NLP Preprocessing Tasks to Be Done for Text Analysis
  • Example Project/Application on Twitter Text Mining and Sentiment Analysis in R

  • Important Notes !!
    You Have to Decide Which Text Cleaning Preprocessing Should Be Applied First in a Pipelining or Should Not Be Applied Depending on Your Method of Text Analysis

    For TF-IDF Based/Lexicon Based Text Vectorization, You Should Apply All the Text Cleaning Preprocessing, Stemming/Lemmarization, and Stopword Removal.
    On the other hand, for the Context Aware Parsers for POS or NER Tagging, or for Training with Sequence of Tokens in a Sentence, Each sentence should be preserved as it is with the order of tokens.
    You Should Not Remove Each Sentence Strcuture, so Do Not Apply These Common NLP Cleaning, StopWord Removing, or Stemming/Lemmarization.



    To Get Lexicon Word(Dictionary) Lists
    WordNet -- Largest Lexical Database of English
    WordNet and Synset Download


    General (Short) NLP English Word List for Simple Text Analysis Applications:

    Stop word List
    Positive word List
    Negative word List


    Sample Big Data Projects with Text Analysis:

    Sample Big Data Project: Building a Content Based Document(wiki Pages) Search Engine System
    Sample Big Data Analytic Project: Building a Intelligent Document Ranking System Using Machine Leraning
    Sharesci Github Site for Code Examples
    Sharesci Wiki Site for Documentation




    Advanced Study on Clustering Analysis Using Similarity Measures for Document Categorization

    Clustering Analysis Concept
    Clustering Algorithms in scikit-learn







    11-12



    Problems (Limitations) with Lexicon Based TF-IDF for Text Analysis:

    Phrase (N-Gram Word) Identification
    Negation Handling
    Can NOT Identify Relationships among Terms - Synonyms (Similar Meaning) or Antonyms (Opposite Meaning) of Terms
    Can NOT Identify Different Meanings of a Same Term by the Different Context of a Senetence - Polysemy Problem !
    Can NOT Identify New Terms, Common Slangs or Changining Relationship Between New Terms (ex: Data Analytics) and Old Terms (ex:Data Mining)
    Can NOT Identify Sarcastic Contexts


    Solutions for the Probelm that TF-IDF Can NOT Identify Different Meanings of a Same Term by the Different Context of a Senetence

    Context Aware Learning with Sentence/Paragraph/Document as it is
    Lectures:

    Context Aware Natural Language Processing(NLP) Methods:

    POS (Part of Speech) Tagging:

    Lecture Notes on POS (Part of Speech) Tagging (Stanford)
    English Grammar
    UPenn Treebank POS Tagging List


    NER (Named Entity) Tagging and Information Extraction (IE) :

    Lecture Notes on NER (Named Entity Recognizer) (Stanford)

    Software NER Tagging Parser:

    Stanford NLP NER (Named Entity) Tagging

    Software Open IE (Information Extraction) Parser:

    Stanford NLP Open IE (Information Extraction) for Triple Extraction
    Github site of Stanford NLP Open IE (Information Extraction) for Triple Extraction


    Useful Natural Language Processing (NLP) APIs to Use and Resource Sites for Download to Use for Text Analysis Applications:

    Core NLP Software Site of Stanford
    Core NLP online Run Test Site (Stanford)
    Avaliable Stanford NLP Softwares





    How to Build a Sentiment/Opinion Analysis Application with Text Analysis Techniques -- For Final Project !!

    Lecture Notes_2_3: Overview of Sentiment Analysis System

    Overview of Text Processing Tasks for Sentiment Analysis of User Review Texts :

    Class Note_18_1: Lecture Notes On Summarizing Opinions of Sentiment Feature (Aspect) Words in Review texts
    Class Note_18_2: Lecture Notes On Research for Sentiment/Aspect Analysis Algorithms/Methods



    Examples of Simple Sentiment Analysis:
    Project Examples of Vectorization of Each Document to Label (Ground Truth) of a Training set for Classification for Sentiment Analysis:

    Project Example I: Sentiment Analysis of Yelp Business Review with Machine Learning from a Research Paper
    Project Example II: Sentiment Analysis of Yelp Business Review with Machine Learning




    For Advanced Sentiment Analysis Proejct:
    2020 Presidential Election Prediction from Social Network Twitter Analysis

    Well Known Ranking Algorithms to Score a Sentiment Score of each Token Word for Vectorization of Document for Labeling (Ground Truth) of a Training Set for Sentiment Analysis:

    1. Senti-WordNet for General Context of Texts:

    Senti-WordNet for Sentiment Analysis
    Presentation of Paper: Senti-WordNet for Sentiment Analysis
    Paper: Senti-WordNet for Sentiment Analysis


    2. Social Media (Twitter) Text Specific Ranking Algorithm:

    Ranking Algorithm: VADER for Sentiment Analysis
    Compound Score of VADER Ranking Algorithm for Sentiment Analysis
    Paper: VADER for Sentiment Analysis






    12-13




    Machine Learning for Data Anaytics

    Lectures:

    Lecture Notes: What is Data Mining with Overview of Data Mining Steps



    Important Data Mining Steps:

    1. Feature Selection for Training Set for Machine Learning


    2. Data Preprocessing Methods for Data Cleaning for Features for Training Data for Machine Learning:

  • Data Cleaning

  • Lecture Notes on Data Properties, Data Structures, and Data Cleaning Issues for Machine Learning


    Tutorials for Data Preprocessing (Cleaning) Methods:

  • Null Value Removal (or Replacement) Methods: Replace Nulls with 0, Min, Mean/Median
  • Outlier Handling Methods
  • Exploring Features with Basic Data Statistics
    What is Outlier?
    IQR Method for Identification and Removal(or Replacement) of Outliers





    3. Data Transformation Methods

    Data Preprocessing Methods for Feature Transformation for Training Data for Machine Learning:
  • Data Transformation

  • Lecture Notes on Data Transformation: Normalization, Discretization

    - DT(Decision Tree) and RF(Random Forest) Require Data Discretization


    - ANN and SVM Reqire Data Transformation:

  • Normalization for Any Numerical Attributes
  • Binarization (One Hot Encoding) for Any Categorical Attributes

  •    Tutorial for Data Preprocessing (Transformation): Normalization and Binarization (One Hot Encoding) for ANN or SVM

    Binarization for Categorical Features :

    Why One Hot Encoding?
    How to Transform Categorical Features
    Data Preprocessing Methods for Numerical and Categorical Features for ANN or SVM


    Transformation of TimpeStamp Features
    Transformation Methods for DateTime Data Type

    Examples of Preprocessing DateTime as Features for Classification: Appointment Cancellation Prediction






    4. Supervised Learning with Machine Learning

    Classification with Machine Learning:

    Concept at First Glance: Binary Classification, Multinomial Classification

    Understanding SVM Algorithm Deeper: How SVM Works for Classification

    Complete List of scikit-learn Supervised Learning Algorithms


    Machine Learning with Probabilistic Approach:

  • Decision Tree:

  • Lecture:
    Lecture Notes on Basic Classification with Decision Tree

    Lecture Notes on CART wth Case Study

    Corrections in Lecture Notes !

    Python Implementation:

    scikit-learn: Classification with Decision Tree


    Random Forest: Advanced Decision Tree with Ensemble Method

    Random Forest (Decision Tree Based)
    Random Forest in scikit-learn




  • Naive Bayes

  • Lecture:
    Lecture Notes on Basic Classification: Naive Bayes

    Python Implementation:
    scikit-learn: Classification with Naive Bayes





    Machine Learning with Numerical Approach:

  • Artificial Neural NetWork
  • SVM (Support Vector Machine)

  • Lecture:

  • Artificial Neural NetWork


  • Lecture Notes on Intro to Classification: Concept of Artificial Neural Network, Support Vector Machine (SVM)

    Advanced Lecture Notes on Artificial Neural Network

    Backpropagation Algorithm with Example


    Step by Step Implementation with Basic Examples for Lectures
    Walk Through with Basic Example using Panda for Classification with Neural Network
    Easy Tutorial on Neural Network Concept


    Basic Underlying Concept of Regression as Classifier:
    Regression Model
    Logistic Regression
    Walk Through with Python Example for Classification with Simple Regression Algorithm


    Easy Tutorial on Multi Layor Neural Network with Softmax Function

    Easy Tutorial on Softmax Function in NN



    Tutorial for Python Implementation:

    Walk Through with Basic Python Example using Panda for Classification with Neural Network

    Walk Through with Basic Python Example in Keras Framework for Classification with Neural Network

    Keras with Tensorflow for ANN

    TensorFlow Example of Classification with Neural Network

    scikit-learn: Classification with Neural Network




  • SVM (Support Vector Machine)

  • Tutorial for Python Implementation:

    Understanding SVM Algorithm Deeper for Classification with Python Example
    Walk Through with Basic Example for Classification with SVM

    scikit-learn: Classification with SVM with Code Examples

    scikit-learn: Classification with SVM

    SVM Kernel Functions





    Example of Application for Intrusion Detection

    Example Codes of Classification Process for Intrusion Detection










    5. Evaluation Metrics for Models of Classification:

    Lecture:
    Lecture Notes on Evaluation of Classification Model

    Cross Validation

    Cross Validation to Avoid Overfitting and Underfitting Probelms

    Bootstrap Sampling Method

    Stratified Sampling Method

    Python Implementation:
    sklearn metrics:f1_score

    evaluation metrics for unbalanced class data set: Macro and Micro f1 scores





    The most important tasks to learn in Classification:

    - Preprocessing methods to transform bigdata to accurate vector form.
    - Understanding the algorithms with the differences in input parameters
    - Understanding the problems, limitations of each ML, and important factors that affect a ML model accuracy and - Solutions to fix the problems and limitations
    For example, Try multiple classifications with SVM and different kernel functions to see which kernel function generates the best model in accuracy.


    Important Problems in Classification

  • Curse of Dimensionality
  • Feature (Selection) Reduction Methods
    - PCA (Principle Compoment Analysis) -- Traditional Statistical Method
    - Correlation Analysis
    - ML Based Recursive Feature Elimination (RFE) -- Advanced ML based Method


  • Model Overfitting and Underfitting Problem


  • Class Imblance (Minority Class) Problem
  • Handling Imbalanced Data for Classification:
    Learn how to do random sampling with SMOTE for Imbalance Data for Anomaly Detection

    Handling Imbalanced Data with SMOTE in Python
    Tutorials with Examples:
    Common Sampling Methods to Handle Imbalalnced Class in Python
    Handling Imbalanced Data with SMOTE in Python







    Data Sets for Classification:

    Kaggle Data Set Repository
    UCI Data Set Repository for Classification
    Health Data Set Repository for Classification
    Socrata Open Data API






    Natural Language Processing (NLP) with Machine Learning: Extra Credit !

    Semi-Supervised Learning for Text Analysis:

    Context Aware NLP Methods Using Machine Learning with Semi-Supervised Learning: Word2Vec (Google)

    Word2Vec:

    Tutorials for Word To Vector: Skipgram Model

    Theory to Understand How Word2Vec works with Skipgram Training Model (Advanced)



    Tutorial Sites for Word2Vec Implementation:

    Tutorials for Word To Vector with Google Tensorflow
    Tutorials for Word Embeddingd with Google Tensorflow


    To get an Executable Binary of word2Vec Model Implementation and Training Data sets:

    Implementation sites of Word2Vec from Google AI 2013:
    Google Original Sites for Word2Vec

    Extended Word2Vec: Glove from Stanford:
    Sites for Glove Word Vector by Stanford NLP Team

    Other Related Sites
    Sites for Word To Vector for npmjs
    CNTK: Stanford NLP Sentiment Analysis



    NLP in Deep Learning Language Model:
    Google BERT (Bi-directional Transfomer)

    What is Google AI BERT Transformer?

    BERT: What this means to your Web Search?

    SpaCy Embeddings for BERT Transformer



    13-14



    Big Data Processing Systems and Parallel Distributed Programming Paradigm


  • MapReduce and Hadoop Distributed File System (HDFS)

  • Map Reduce by Google:

    MapReduce Lecture:
    Class Note_18: Lecture Notes On Map Reduce by Google

    Exampel of Unstructured Text data Processing for HDFS with MapReduce:
    Class Note_18_1: Inverted Index in Google (Stanford)

    Original Research Paper: Map Reduce by Google



    REVIEW on RDBMS File System:
    Bsics of RDBMS File System

    Lecture for Hadoop (HDFS) by Apathe:

    Class Note_18_1: Introduction to Hadoop Distributed File System (HDFS) (Introduction Only)
    Code Example of Word Count MapReduce in Hadoop Version 2.0


    Hadoop Job Class

    MapReduce Tutorial:
    Map Reduce Tutorial on Hadoop


    Hadoop Implementation Set Up:
    Hadoop Single Node Set Up
    Hadoop Fully Distributed Mode (Cluster) Set Up


    Code Examples of MapReduce on Hadoop:
    Simple MR Hadoop Lab Guides
    MR sample Codes - Average Temp By Zip Code








  • Spark: In Memory Processing for Streaming

  • Class Note_26_1: Lecture Notes On Spark: Spark Streaming/SparkSQL/SparkR

    PySpark Platform :

    PySpark for Beginners
    PySpark Tutorial for Beginners
    PySpark Tutorial
    PySpark Tutorial for Data Science


    15


    Cloud Computing

    Class Note_21: Large Scale Big Data WebApps I
    Class Note_22: Large Scale Big Data WebApps II

    Multimedia Service on Cloud
    Mobile, Multimedia and Cloud Computing
    Class Note_11:Lecture Note on Cloud Computing


    Amazon Cloud Service:

    Amazon AWS RDS (Relational Database Service)
    Amazon AWS Elastic Cloud Computing (EC2) for Web service
    How to Create Amazon AWS Virtual Hosting (EC2) with a Web Server and Amazon Database Server (RDS)

    Microsoft Cloud Service Azure

    Trial account for Microsoft AZURE Cloud
    AZURE Cloud
    Tutorial for MS AZURE Cloud
    Sample Project on Microsoft AZURE Cloud
    How to Create/Retrieve a Table in Microsoft AZURE Cloud
    How to Create/Retrieve BLOB data in Microsoft AZURE Cloud
    How to Create a SQL Database Server in Microsoft AZURE Cloud


    6-7


  • Big Data Platform in Web Applications

    Web Server Programming with Node JS and Mongo DB (MEAN Stack)


    Serverside JavaScript Framework: NODE JS

    Class Note_18: Controller Server Communication
    Class Note_19: Introduction to Web Server
    Class Note_20: Introduction to Node JS
    Class Note_21: Express

    Good Node JS/Express Tutorial Site
    Express Basic Routing
    Angular route Parameters



    NodeJS Set Up Guide

    NodeJS Site
    NodeJS Download
    NodeJS Package Manager Download
    NodeJS Installation Tutorial on Window
    NodeJS Installation Tutorial on MacOS
    NodeJS Installation Tutorial on Ubuntu/Linux

    Another NodeJS new source site for NodeJS Set Up
    NodeJS Download



    Node JS Framework: Express
    Node JS Framework: Express
    Express Set Up: Getting Started

    Express Coding Guide:
    Express Example: How to Create a Simple App
    Express Tool to Generate an Express App
    Express Basic Routing
    Angular route Parameters
    Express Guide: Routing to Write
    Node JS HTTP Methods to Write

    Basic Express Built In Frameworks/Middleware:
    Node JS Documentation: Global variables
    Node JS Documentation: _dirname, _filename
    Express Guide: API Methods: res.json()
    Express Guide: Writing a Server Side Code:Middleware
    Express Guide: Integrating a Database
    Express Guide: Static File Server
    Node JS: server.address()
    How to Know your server address on Window
    Node JS: multer for file upload
    Node JS: bodyparser.urlencoded with extended false/true
    Node JS: bodyparser.urlencoded with qs
    Node JS: bodyparser.urlencoded

    Node JS: request object
    Node JS: Node JS: response object
    Node JS: setting response object
    Node JS: http API



    Other Resourses:
    Node JS Documentation
    Framework Built on Node JS Express





    NODEJS SAMPLE APPLICATION CODES

    Web Server Code Examples with NodeJS and Express :

    Good Node JS Express Tutorial Site

    Tutorial for Express Basic Routing

    First Node App






    Integrating a Database Server with Node JS: MySQL/MongoDB/PostgreSQL/SQL Server:

    Class Note_22: Database Server

    Node API for MySQL:
    NodeJS: Integrating to MySQL
    Node-MySQL


    NodeJS works with Most of Databases:
    Example Codes for Connectivity to Each Database in NodeJS

    Database Connectivity in NodeJS

    Node Package Manager (NPM) for Database Connectivity

    Microsoft SQL Server: Database Connectivity in NodeJS
    MySql: Database Connectivity in NodeJS
    Postgre: Database Connectivity in NodeJS



    Application Code Examples with NodeJS with Database Server MongDB or MySql :


    Simple Node Application Codes with MongoDB

    Node Application Codes in MVC with MongoDB

    Node Application Codes in MVC with MySql





    Set up Guide for MEAN Stack: Also See Project 3 Section For Node JS Set up Guide
    Node JS with Mongo DB Setup Guide
    Node JS with Mongo DB Setup Guide
    Angular JS with Node JS with Setup Guide

    Node JS API for MongoDB:
    NodeJS: Integrating to MongoDB
    Mongoose: Node JS API for MongoDB: For connecting/querying to MongoDB:
    Mongoose To Set Up
    Mongoose Guide


    Application Examples Built with Node JS
    Sample Web Application Using Node JS with Mongo DB
    Sample Web Application with Angular JS and MS SQL Server



  • 16

    Project Presentation
    Project Presentation Scedule

    ==> Completion of Homeworks/Labs is required for obtaining a passing grade.

  • This is a tentative scale and
    it could be changed

    Letter
    Grade

    Quality Points

     


      A

    > 93%  

    A: Outstanding (student's performance is genuinely excellent)

      A-

    90% - 93%

     

      B+

    87% - 90%   

     

      B

    82% - 87%

    B: Very Good (student's performance is clearly commendable but not necessarily outstanding)

     

      B-

    80% - 82%

     

     

      C

    75% - 80%

    C: Good (student's performance meets every course requirement and is acceptable; not distinguished)
        D 65%-75% D: Below Average (student's performance fails to meet course objectives and standards)

     

      F

    <65%

    F: Failure (student's performance is unacceptable)

    ADA Adherence. If you need course adaptations or accommodations because of a disability, if you have emergency medical information to share with me, or if you need special arrangements in case the building must be evacuated, please make an appointment with me as soon as possible. My office location and hours are listed on top of this syllabus. If you need further information, please contact the ACCESS office, phone number 687-5106.

     


    Programming standards

    • Every program must include your name, CSU ID number, Class, Section Number, Hours, the words 'Homework # ...', and a short description of the assignment. For example:
       ' Name: Mark Zuckerberg  
       ' ID: 1234567            
       ' Homework #1            
       ' Description: Computing the average life of a light bulb
    • Every variable should have a meaningful name (this includes function/procedure/subprogram names).
    • Every portion of the program should be as cohesive (single purposed) as possible. This leads to a large number of small functions.
    • Every function (including the main function) should be preceded by a comment indicating its arguments and a description of the transformation it performs.
    • Non-obvious code within a function should be explained.
    • Code should not be over commented.