CIS 660 Data Mining with Advanced Machine Learning (4-0-4) |
||
| Course Content |
|
|
|
Semester Schedule: //a href="https://www.csuohio.edu/registrar/academic-calendar"> See University's Official Academic Calendar for the Semester Schedule to Add and Drop and the Final Exam Schedules 19. August 25, 2025: The CIS660 Final Exam Info: Final Exam schedule for Fall 2025: Wed Dec 10, 12:30PM - 2:30PM in Class For the Active Lecture Notes Links for the CIS660 Final Only, Access the following site: The rest of links are all disabled. //a href="https://eecs.csuohio.edu/~sschung/CIS660/CIS660Fall25_FinalOnlyLinks.html"> the Active Lecture Notes Links Only for the CIS660 Final Only One Page Note is Allowed to the Final. The Exam Format and Rules will be the same as the Midterms (Short Exams). Subjects to Focus On for the Final Important Problems to Resolve to Increase Model Accuracy of Classification Clustering Algorithms - K-Mean, Heirarchical Clustering Algorithms, DBScan, Accuracy Estimation and Validation, Model Evaluation Methods, Model Comparison Methods (NOT for this Semester) Advanced NLP: Word2Vec, BERT RNN, LSTM Deep Learning Architecture with CNN (Not for This Semeser): Association Rule Mining: Apriori Algorithm and Optimization, Assocaition Rule Generation, Frequent Pattern Tree, Interesting Measures 18. August 25, 2025: Midterm Info: //a href="CIS660_MidtermOnly_Links.html"> Midterm Info and The Lecture Note Links Only for Midterm There Will Be 2 Exams (15% Each) Subjects to Focus on: Chapter 2, 3 on : Basic Stats of Data, All the Data Proximity Measures, Data Transformation Methods, Feature Selection Methods, Feature Correlation Measures -- Kai Square, Correlation, Covariance Basic Text Analysis Algorithms and Methods in Information Retrieval: TF-IDF Measure, Document as Term Vector Note that You are responsible ONLY for the Lecture Notes and the Subjects Covered in Class Question Types: 6-7 Short Answer Questions on Problem Solving with Given Small Data Sets No T/F Questions, No Multiple Choice Subjects to Focus on: The Second Exam on Oct 30 on Machine Learning Algorithms. The Subjects Covered on Machine Learning Algorithms and Their Objective Functions Covered in Class by Mon Oct 28: All the Lectures on Decision Tree, K-NN, Naive Bayes, Baysian Belief Network, Ensemble Methods Linear Regression, Logistic Regression, Multinomial Logistic Regression, ANN Feed Forward, Backpropagation Algorithms, Ensemble Methods -- Bagging, Boost, Ada Boost, Sampling Techniques, Model Evaluation Methods Common Problems in Classification and ML Models, Data Preprocessing Methods of Each Classifier The Same Question Types as the Exam 1 15. August 25, 2025: There Will be Random Quizzes to Check the Attendance ! Only 5 - 10 Mins Will Be Allowed for Each Quiz. Those Who Come to the Class Late, They Will NOT Be Given a Quiz ! 13. August 25, 2025: Only the registered students can access the course blackboard. If you have a problem with your blackboard access, please contact the registrar or CSU Tech Support to resolve the issue ! //a href="https://www.csuohio.edu/center-for-elearning/technical-support"> Center-for-elearning Technical Support Faculty does not control your registration and the course blackboard access in the CSU systems. 3. August 25, 2025: TA Information: TA: Email: j.guo58@vikes.csuohio.edu Office Hours: Mon and Wed 4:00 pm -6:00 pm (Send him email ahead to let him know you are coming) Location: FH311 or ZOOM Meeting br> ZOOM Meeting ID: Passcode: //a href=""> If you have questions in Labs or Grading your Labs, Send an email to TA or Talk to him During his TA Office Hours or Schedule a Zoom meeting with him Dr. Chung's Office Hours: Tues and Thursday 1:30PM - 3:30PM Office: FH 222 or Zoom Meeting Email: s.chung@csuohio.edu Send Me Email to Set Up a Meeting or Zoom Meeting. Meeting ID: 859 3867 6332 //a href="https://csuohio.zoom.us/j/85938676332"> Zoom Meeting Link 2. August 25, 2025: The Output of each lab is your report in Doc file that shows your screen captures of your data processing steps for analytics and the results Each of your screen capture must show each step and the result returned by the server in the SAME window in your System to prove that your lab is done correctly in YOUR SYSTEM !! Lab Submission: 1. Submit your Lab in zip file including 1) your lab report in .doc and 2) all the source files, preprocessed input files, outputs on Blackboard for a timestamp as a proof. 2. If You did an Extra Credit Part, Mention about What Part is Done for Extra Credit at the Heading of the Front Page of Your Report in Bigger and Bold Font ! 1. August 25, 2025: The class webpage link is announced in the Class Blackboard as well. //a href="https://eecs.csuohio.edu/~sschung/CIS660/CIS660Fall25.html"> Class Webpage Semester Schedule: //a href="https://www.csuohio.edu/registrar/academic-calendar"> See University's Official Academic Calendar for the Semester and the Final Exam Schedules Final Exam schedule for Fall 2025: Wed Dec 10, 12:30PM - 2:30PM in Class Midterm Info: //a href="CIS660_MidtermOnly_Links.html"> Midterm Info and The Lecture Note Links Only for Midterm There Will Be 2 Exams (15% Each) Subjects to Focus on: Chapter 2, 3 on : Basic Stats of Data, All the Data Proximity Measures, Data Transformation Methods, Feature Selection Methods, Feature Correlation Measures -- Kai Square, Correlation, Covariance Basic Text Analysis Algorithms and Methods in Information Retrieval: TF-IDF Measure, Document as Term Vector Note that You are responsible ONLY for the Lecture Notes and the Subjects Covered in Class Question Types: 6-7 Short Answer Questions on Problem Solving with Given Small Data Sets No T/F Questions, No Multiple Choice Subjects to Focus on: The Second Exam on Oct 30 on Machine Learning Algorithms. The Subjects Covered on Machine Learning Algorithms and Their Objective Functions Covered in Class by Mon Oct 28: All the Lectures on Decision Tree, K-NN, Naive Bayes, Baysian Belief Network, Ensemble Methods Linear Regression, Logistic Regression, Multinomial Logistic Regression, ANN Feed Forward, Backpropagation Algorithms, Model Evaluation Methods Common Problems in Classification and ML Models, Data Preprocessing Methods of Each Classifier The Same Question Types as the Exam 1 Important Note for Exams: I am afraid that it is not possible to change the schedule of the midterm or Final exam for one person's favor. |
The prerequisite of CIS660 has been changed to CIS530 and CIS550, which was intended to prevent the students without any CS/DS Undergrad Background from Registering for CIS660 in their first semester without completing preparatory courses.
CIS530 and CIS550,Engineering Statistics are Required, and an Undergrad (Intro) of Machine Learning or Algorithms Are Preferred as Prerequistes of CIS660.
|
Labs: First Week Lab0: Choose your System/Tool/Platform to Set Up and Get Used to: See the Lab Assignment0 Section Below to See the Set Up Guide for Python Data Science Platform in Step by Step ! Basic Python Tutorials : Python Data Science Platforms There Are Mainly Three Ways to Set up All the Neccessary Data Science Software/Library below ! See the Instructions to See Three Options and How to Set Up the Debugger Spider Here !! //a href="https://eecs.csuohio.edu/~sschung/DSA460/DataSciencePlatform_Setup_Guide.pdf"> Set Up Guide for Python Data Science Platform in Step by Step ************************************ Additional Installation Guide: //a href="https://www.scipy.org/install.html"> Installation Guide for Scientific Python tools for Data Science with pip (inbuilt package management system) Python IDE Deduggers: //a href="Pycharm_Debugger"> Basic Guide for Python Debugger Pycharm //a href="https://www.spyder-ide.org/"> Python IDE Spyder //a href="http://docs.spyder-ide.org/current/panes/debugging.html"> Python Debugger Spyder Python Data Science Platforms: //a href="https://www.scipy.org/install.html"> Installation Guide for Scientific Python tools for Data Science with pip (inbuilt package management system) //a href="https://www.dataquest.io/blog/jupyter-notebook-tutorial/"> Basic Guide for Python tools for Data Science: jupyter-notebook //a href="https://www.dataquest.io/blog/advanced-jupyter-notebooks-tutorial/"> More Basic Guide for Python tools for Data Science: jupyter-notebook • Anaconda //a href="https://www.anaconda.com/open-source"> Anaconda Open Source Site See Fundamental Section for List of Data Science Platforms //a href="https://docs.anaconda.com/anaconda/navigator/tutorials/"> Anaconda Tutorials • Python Anaconda Tutorial Sites //a href="https://data-flair.training/blogs/python-anaconda-tutorial/"> Anaconda Tutorial Site //a href="https://www.edureka.co/blog/python-anaconda-tutorial/"> Anaconda Tutorial Site • PyTorch //a href="https://pytorch.org/"> PyTorch Site (It can be integrated from Anaconda as well) Google Colab: • Python Scikit Learn for Common Data Science Tasks Text Preprocessing (Natural Language Processing) Library in Python SpaCy: //a href="https://spacy.io/api"> Liquistic Modules in Python SpaCy //a href="https://spacy.io/api/lemmatizer"> Lemmatizer in Python SpaCy //a href="https://stackabuse.com/python-for-nlp-tokenization-stemming-and-lemmatization-with-spacy-library/"> Liquistic Modules for Tokenization, Stemming, Lemmatization in Python SpaCy //a href="https://stackoverflow.com/questions/38763007/how-to-use-spacy-lemmatizer-to-get-a-word-into-basic-form"> How to Code Liquistic Modules like Lemmatizer in Python SpaCy //a href="https://codeburst.io/python-basics-11-word-count-filter-out-punctuation-dictionary-manipulation-and-sorting-lists-3f6c55420855"> Python Example for Basic Text Processing • Python sklearn.cluster //a href="https://scikit-learn.org/stable/modules/clustering.html"> Python Sklearn Clustering There is Another Data Science Platform in R if You Choose to Learn (SAS Data Miner or MS Data Tool Have an Integrated R Platform) Basic R Tutorials : //a href="IndependentStudyCIS611Final Report.pdf"> Special Online Study Guides on Basics on Data Warehouse/OLAP, Data Analytics, Big Data in Independent Study by Nick White (Now in FaceBook) Useful Machine Learning Tutorial Sites: //a href="https://keras.io/guides/"> Keras for Image Processing/Text Processing with Deep Learning //a href="https://colab.research.google.com/"> Google Colab for FAST Machine Learning Execution in GPU Good Data Science and AI Platforms for Developing and Training: For Your Own Study Lab Submission Instructions: The Output of each lab is your report in Doc file that shows your screen captures of your system/tool/platform setting/configurations, your data processing steps for analytics and the results Each of your screen capture must show each step and the result returned by the server in the SAME window in your System to prove that your lab is done correctly in YOUR SYSTEM !! 1. Submit (on Blackboard) your Lab in a zip file including 1) your lab report in .doc and 2) all the source/scipt files, preprocessed input files, outputs on Blackboard for a timestamp as a proof. 2. If you did an Extra Credit Lab, Make a Note on the COVER of Your Lab Clearly ! See the Instructions on How to Create Your Lab Report Below ! The Lab Submission Link and the Deadline of Each Lab Will Be Posted on the Class BlackBoard ! You Have to Start Working on Labs Before the Submission Link Are Created on Blackboard for Each Lab Submission If You did an Extra Credit Part, Mention about What Part is Done for Extra Credit at the Front(Cover) Page of Your Report in Bigger and Bold Font ! Always Follow the Deadline of Each Lab Assigned on the Class Blackboard. The Deadlines mentioned on the Class Webpage Are Tentatively Scheduled at the Beginning of Each Semester. Please Identify Your Course When You Ask Me in Email ! 1. Submit your Zip file that includes your report in .doc file (that expains all the platform set up, the execution steps, and copy of each source code files ) and all the Source files, and output files on Blackboard for a timestamp and as a proof. 2. IMPORTANT NOTE: Your Screen Captures in Your Lab Report Should Show Your Own System Name to Prove That Your Lab Was Done In Your System. 3. If You did an Extra Credit Part, Mention about What Part is Done for Extra Credit at the Front(Cover) Page of Your Report in Bigger and Bold Font ! 4. Always Follow the Deadline of Each Lab Assigned on the Class Blackboard. The Deadlines mentioned on the Class Webpage Are Tentatively Scheduled at the Beginning of Each Semester. 1. Submit your Zip file that includes your report in .doc file (that expains all the platform set up, the execution steps, and copy of each source code files ) and all the Source files, and output files on Blackboard for a timestamp and as a proof. 2. IMPORTANT NOTE: Your Screen Captures in Your Lab Report Should Show Your Own System Name to Prove That Your Lab Was Done on Your Computer. 3. If You did an Extra Credit Part, Mention about What Part is Done for Extra Credit at the Front(Cover) Page of Your Report in Bigger and Bold Font ! 4. Always Follow the Deadline of Each Lab Assigned on the Class Blackboard. The Deadlines mentioned on the Class Webpage Are Tentatively Scheduled at the Beginning of Each Semester. Lab Assignment 0: Set Up Your Python Data Science Platform for Labs and Final Projects by the End of the First Weekend ! There Are Mainly Three Ways to Set up All the Neccessary Data Science Software/Library ! See the Instructions to See Three Options and How to Set Up the Debugger Spider Here !! Optional for Big data Processing Although the Database Skills and Knowledge are not required for this course, it will be very useful for handling big data sets. For Those Who want to Use a Database Server for Initial Data Handling //a href="https://eecs.csuohio.edu/~sschung/cis430/CIS430IDS.html#Lab"> See the Lab Sections of CIS 430/530 for SQL Server Installation Guides //a href="https://eecs.csuohio.edu/~sschung/cis611/CIS611IDS.html#Lab"> See the Lab Sections of CIS 611 for SQL Server/Data Warehouse OLAP Server Installation Guides REQUIRED LAB ASSIGNMENTS: The Submission Link and the Deadline of Each Lab Will Be Posted on the Class BlackBoard ! Lab Assingment 1: Lab1_1 - Part 1 and Part 2 on Preprocessing and Transformation: Due By the End of the Third Week Lab1_2 - Part 3 and part 4 on Similarity Measure and Correlation Matrix: Due By the End of the Fourth Week You don't Have to Import AdventureWork Data Warehouse to Get the View vTargetMailCustomer for Lab1. You can directly download vTargetMailCustomer.csv below. Input Data File for Lab1: For Part 1 and 2: For each selected feature, identify ALL the required data preprocessing methods like Normalization, Discretization, Binarization, and more based on the feature data properties. They must be done along with other preprocessing methods. Examples of Lab1 Report of Data Preprocessing and Data Similarity Measures //a href="Lab1_OutputExample_1.pdf"> Example1 of Lab1 Output: Part 1 //a href="Lab1_OutputExample_Part3_Similarity.pdf"> Example of Lab1 Output: Similarity Measure Note that these Examples Here May NOT neccessarily all correct ! They just show How Lab1 can be done as an example. Do NOT blindly follow ! For example, EnglishEducation should be Transformed as Categorical or Ordinal? Depending on the properties of the column, the corresponding transsformation method should be applied. For Part 3 and 4: IMPORTANT NOTES !!! You Are Not Supposed to Use any Python Libs to Calculate Similarity Measures to Build a Similarity Matrix or Correlation to Build a Correlation Matrix in Part 3 and Part 4 of Lab 1. You have to write Scripts/Programs to Compute each Similarity or Correlation and Build a Similarity Matrix or Correlation Matrix. If you use the built-in lib to build a Correlation Matrix, which is one line of code, you will get 0 for the part. Suggested Platforms to Use for Lab1: Anaconda //a href="https://www.anaconda.com/open-source"> Anaconda Open Source Site See Fundamental Section for List of Data Science Platforms //a href="https://docs.anaconda.com/anaconda/navigator/tutorials/"> Anaconda Tutorials Lab Assingment 2: You can use any avaialble API such as Beautiful Soup for cleaning webpages in html tags Some Trouble Shooting Tips: For Lab 2, it requires the python library 'clean-text' The full command to install this dependency is: pip install clean-text If you do wish to rerun it and make sure that it works, you would have to uninstall the old one first as they use the same module name. Having both installed at the same time will result in Python using the incorrect one. NOTE !!! You Are Not Supposed to Use any Python Libs to Build a Cosine Similarity Matrix for Lab 2. You have to write a script to compute and build a Similarity Matrix. If you use the built-in lib to build a Cosine Similarity Matrix, you will get 0 for Labs. Lab Assingment 3 on Classification IMPORTANT NOTES for LAB3: 1. Use the Best Selected Features Given Below instead of Your Own Selected Feature Set in Lab1 : IMPORTANT !!: In your Lab report, Make sure to Show the final transformed training set file and all the attribute values for the first two objects in your training set that was used for your classifiers. 2. Choose your Classifiers for Lab3 as below: 1) One from Decision Tree or Bayesian and 2) Another from Any Ensemble Methods: Random Forest, XGBoost, LightGBM, Adaboost 3) ANN 4) SVM or 4) One from Distance Based K-NN 3. Grading is Based on Your Best Accuracy with the best input hyperparameters you identified from Your Experiments Lab Assingment 4 on Clustering: 2. Experiment to find the best Input Parameters for each Algorithm. 3. For each Clustering result in your experiment, Apply any method discussed in the Lecture notes (Ward's method, Silhouette score, Elbow method, Entrophy/Purity) to Measure the quality of the Clustering result. 4. Visualize the final best clusters Data Sets you can Choose for Lab4: Some Suggested Platforms for Clustering //a href="https://scikit-learn.org/stable/modules/clustering.html"> scikit-learn Clustering Although Clustering in general should work on data as multidimensional vectors, For NIJ Data For Lab4, you can work on NIJ Challeneges to identify Hot Spots for Crime Location - GPS data and Crime Category instead of handling data as multidimensional object for Clustering algorithms Cluster on the Crime Loactions (X and Y coordinates - transform them to GPS coordinates) with K-Mean and DBCSAN and Cluster the Crime locations for each Crime Category but Add PAI or PEI Analysis for Hot Spot Analysis with Changing Parameters National Institute of Justice (NIJ) Crime HotSpot Analysis Challenging: NIJ Crime Location Data Sets: AWID data Set: EXTRA CREDIT (see more info in the Clustering and Anomaly Detection Section) AWID Site: //a href="https://icsdweb.aegean.gr/awid"> AWID: Intrusion Detection in Wireless Network Server Log Data Apply a Feature Selection Method Discussed in class to reduce the dimensions to apply clustering Suggest to Use EmEditor to open a big file if WordPad can't open it. Sometimes a data file collected from a different file system (HDFS, for example) has incompatible special characters for line feed or some others to open it in a Window/Linux system, so you have to write and run a simple script to replace those special characters in the file. See a sql script to import the AWID to Sql server |
|
Project: Important Dates for Final Group Project Submit Your Group Proposal (in Minimum 3 Pages) on BlackBoard Group Proposal Should Include: 1. Data Description, Data Size, Data Collection Plan (if needed), 2. Systems/Tools to Use, 3. Data Preprocessing Methods, 4. Data Analytic Goal with Plan of Evaluation (Design of Your Experiment) in Detail Your Project Status Report Should Show the Following Tasks Done: 1. Platform Setting/Configuration Procedure if it is new, 2. Your Data Contents, Selected Feature Description, 3. Data Preprocessing Steps and the intermediate Outputs Group Project Presentation Should Include: 1. Data Description, Data Size, Data Collection Method(if needed), 2. Systems/Tools to Used, 3. Feature Selection Method and The Final Features Selected 4. Data Preprocessing Methods and Intermediate Results, Final Training Set and Test Set Description, 5. Data Analytic Goal with (Design of Your Experiment) in Detail 6. The Problems/Errors Encountered and Your Resolutions 7. Evaluation Results and Visualization of the Results Any Group Size in 1 - 4 Person Group Are Allowed Note that Large Groups (3-4 Person Group) Shoud Complete a Bigger Project ! One Submission Per Group Is Required ! EACH Memeber Name and ID SHOULD BE LISTED in the COVER ! List Your First and Last Name Only ! Exactly as Appeared in the CSU CampusNet. DO NOT Use Your Middle Name. If We Can't Find You by Your Name appeared on Your Project Report and Presentation Schedule. Your Project will be Considered as Missing with 0. Task 4: Final Group Project Report Submission Instructions (By the End of Friday of Your Presentation Week): Submit Group Project Presentation and Final Report in a Zip File By the End of Friday of Your Presentation Week ! Remember you have to include the source file of your Project Report in doc and Presentation slides in pptx ! If your data file is too big to upload, Submit your zip file with your Data file on your ONE drive and Send email to me and TA attached the data file on One Drive. Submit a Zip file that includes: 1. All of your presentation slides (both in .ppt and .pdf) and 2. Your Group Final Project Report (in doc) that shows with: 1.Platform/System Set up Procedures/Instructions Make Sure to List the Specific Version of Each Component of Your Platform (Such As the Versions of Python, GPU, Pretrained Deep Learning ML Model, and more) to Run Your Project. 2. Description of the Analytic Goal of Your Project and Description of the Data Set 2. I Need to See Your Transformed Features for each ML Algorithms and Your Source Scripts/Codes. 3. Evaluation Results and Visualization of the Results 4. Executions Steps, all the source codes/scripts, all the intermediate outputs, and final output files 5. Include the Problems/Error Encountered and Your Resolutions in Your Report All These Are Required Becasue I Need to See Proof/Evidence Showing that Your Project Is Not a Copy of a Github Codes You Downloaded from the Internet. One Submission Per Group Required. Submit a Zip File that Includes All the required Source Files, Input, Output Files, and Final Report (in Doc file) and Presentation Slides (in pptx). Your Final Project Report Should Include the Set Up Procedure /Configuration Detail of Your Platform/System/Packages as well as Source Codes and Intermediate Results in files. The Report Should Explain Each Step of Your Project Tasks with the Screen Captures and Results. If you don't show/include any of the required contents in your report and presentation, I will ASSUME that your group submitted a Copy of Somebody's Github Codes from the Web. If you need high computing power system with GPU for training: 1. Good Platforms for Training to Develop Deep Learning Models or Large Language Models: Free GPU Use 2. Big Data Servers with GPU are Available to Use in the Big Data Lab: (Temporary Permission will be given per request) 3. High Computing System in the Engineering College: //a href="https://cis.csuohio.edu/~h.yu/mri.html"> Good Data Sets for Classification: //a href="https://healthdata.gov/stories/s/nqx6-g6vz"> Health Data Sets for Challenges **************** //a href="https://data.gov/"> US Government Data Sets **************** More Project List and Data Sets: Sentiment Analysis of Online Reviews/Social Media data: Research Direction/Guides for Project: Question Answering System/Smart/Intelligent System on Texts/Papers/Documents Recent Research Papers on Large Language Model Training Methods : Prompting Based Training: //a href="https://thegradient.pub/prompting/">Research on Prompting Based Training Recent Papers: //a href="https://arxiv.org/pdf/2210.10341.pdf"> BIOGPT //a href="https://arxiv.org/pdf/2304.03262.pdf"> ChatGPT 3.5 Natively Performs Chain of Thoughts 2023 //a href="https://arxiv.org/pdf/2205.10625.pdf"> Least to Most Prompting 2023 //a href="https://browse.arxiv.org/pdf/2309.15088.pdf"> RankVicuna: Reranking GPT Prompt Model 2023 //a href="TopicSummarizationWeakSupervisionCMU_2020.pdf"> Paper Topic Summarization Weak Supervision CMU 2020 Data Set for Question Answering Training //a href="https://pubmed.ncbi.nlm.nih.gov/download/"> Pubmed Paper Repository Data Sets //a href="https://huggingface.co/datasets/pubmed_qa"> Pubmed QA Data Set at HuggingFace //a href="https://github.com/pubmedqa/pubmedqa"> Pubmed QA //a href="https://github.com/niderhoff/nlp-datasets"> NLP Data Sets How to Collect Social Media Data: (take CIS612 for More on Big Data Processing and NoSQL Big Data Systems) This can be done by collecting Twitter real time stream and apply the basic NLP and Text analysis methods for Sentiment Analysis. You have only about 10 days to collect the data. You have to apply the developer's account to teh Twitter to start. However, it requires some knowledge and skills to collect and process the big data stream in the Twitter logging strucrure before you start any text analytic tasks, which are not the focus of CIS660 because of the extremely limited time to cover all the important analytic subjects. I suggest you to take CIS612 for those subjects. Most of CIS Master/PhD students take both CIS612 and CIS660. CIS 612 (or CIS 593 Big Data) cover all the related subjects for the social media opinion analysis for: 1) How to collect the real time big data like social media Twitter and 2) how to process and manage the collected big data using semi structured database server for further data analytic process. You can try to start collecting the data. Look at the instructions and step by step tutorials in Lab3 section in CIS593 website below. //a href="https://eecs.csuohio.edu/~sschung/CIS593/CIS593IDS.html#Lab"> Lab Section of CIS 593 Big data for Twitter Data Collection (Scroll down to the Lab3 Section) //a href="TwitterDataCollectionMongoDBTableau.pdf">Tutorial: How to Collection Twitter Messages to Process with MongoDB and Tableau //a href="TutorialHowtoGetFaceBookGraphAPIDataFacePager.pdf">Tutorial: How to Get FaceBook Graph API Data to Analytics with Hive General Task Guideline for Sentiment Analysis of Social Media Texts //a href="https://eecs.csuohio.edu/~sschung/CIS660/HowtoDoSentimentAnalysis.pdf"> General Steps for Sentiment Analysis Sample Projects on Sentiment Analysis //a href="ProjectExample_SentimentAnalysis_YelpReview.pdf"> Basic Sentiment Analysis with Yelp Review Data //a href="ProjectExample_SentimentAnalysis_YelpReview_Feng.pdf"> Basic Another Sentiment Analysis with Yelp Review Data //a href="ProjectExample_LDA_Word2Vec_TripAdvisorHotelReview.pdf">ProjectExample on LDA for Topic Discovery and Word2Vec for Similar Words with Trip Advisor Hotel Review Data Lectures to Learn Methods and Find Related Research Papers: //a href="https://eecs.csuohio.edu/~sschung/CIS660/SlidesMiningSummarizingFeatures_HuKDD04.pdf">Paper Presentation on Mining and Summarizing Feature Reviews (KDD 2004) //a href="https://eecs.csuohio.edu/~sschung/CIS660/AspectAnalysisLiteratureSummaryBingLiu2017.pdf">Research Summary on Product Feature Reviews by Dr. Bing Liu //a href="https://web.stanford.edu/class/cs124/lec/sentiment2018.pdf">Research Summary on Sentiment Analysis (Standford) Sentiment Analysis Project based on Earlier Papers: //a href="Sentiment_Pang_2002.pdf">Research Paper: Sentiment Analysis Using Classification //a href="SentimentAnalysis_Pres_Pang_2003Paper_XioadonLiu.pdf">Research Paper Presentation on Sentiment Analysis Using Classification (by Xiaodan Liu) //a href="SentimentPingConelle_2004.pdf">Research Paper: Sentiment Analysis Using Subjectivity Build a Ranking system or Recommendaton System on Product Reviews or Using Real Estate Land Use Data Sets: Using the Real Estate Land Use Data Sets and Working People Demographics in the Cleveland area (Data sets will be given) Generate a Ranking or Recommendation List for a Given User Interest or Prefernece on Real Estate Land Use Data Sets See More Info and Research Papers of Recommendation Systems/Ranking system in the Lecture Notes Section on Recommendation Systems After the Classification Section Project for Recommendation System with Real Estate Land Use Data Set in Cuyahoga County Meta data description is gievn below and the detailed instructions for the access to Data sets will be given later. Permission to access this folder below will be given per request if your group want to do a Project on this data Implement Your Own (Deep) Neural Network Architecture Using Google Tensorflow or MS CNTK or any of your choice for Text Analysis Task or Image Processing: Extra Credit ! Extra Credit for Those Who Took CIS612 Big Data and HDFS: Do Your NN Training in Parallel on Hadoop. Research and learn how to train NN in parallel on HDFS. How to Combine each local model to one final one. If you want to work on Neural Network, then download source codes of skipgram (the first paper) from one of those sites below to learn how to build Skipgarm with NN then implement paragraph vector as document vector in the second paper. Look at the first research project for guide for this project at: Then, let's read the next two papers to understand the source codes for you to download and do the experiment with real data set. Then check these two sites where you can download all the source codes to start an experiment with real data sets. Data set to train to generate word2Vec and Paragraph Vector: You can choose any webpage set or papers as data but it should be at least 200,000 documents to train There is a preprocessed wiki page data set is available in the Stanford NLP site that can be used as your training data set. You can make this as your project if you want. The Most Recent and Most Superior Word Vector: BERT -- See Advanced Text Analysis Section in Class Lecture Notes for more on BERT //a href="https://ai.googleblog.com/2018/11/open-sourcing-bert-state-of-art-pre.html"> BERT from Google AI //a href="http://jalammar.github.io/illustrated-bert/"> Tutorials on BERT from Google Good Tutorial Sites to Learn How to Use Pretrained BERT: //a href="https://spacy.io/usage/embeddings-transformers"> SpaCy Embeddings for BERT Transformer //a href="https://colab.research.google.com/drive/1yFphU6PW9Uo6lmDly_ud9a6c4RCYlwdX"> Good Tutorial Site to Use Pretrained BERT Transformer //a href="https://www.youtube.com/watch?v=l8ZYCvgGu0o&ab_channel=ChrisMcCormickAI"> Good Tutorial to Use Pretrained BERT Transformer for Question Answering System //a href="https://colab.research.google.com/drive/19loLGUDjxGKy4ulZJ1m3hALq2ozNyEGe"> Good Tutorial to Use Pretrained BERT Transformer for Domain Specific System //a href="https://colab.research.google.com/drive/1pTuQhug6Dhl9XalKB0zUGf4FIdYFlpcX"> Good Tutorial Site to Use BERT Transformer for Sentence Classification //a href="PassageRerankingBert2019.pdf"> Research Paper: Passage Reranking Using BERT from Google Anomaly Detection with AWID Data set: Implement one of methods in the recent papers on Anomaly Detection. See More Help on AWID data processing in Lab4 Section ! AWID Site: //a href="http://icsdweb.aegean.gr/awid/features.html">AWID: Intrusion Detection in Wireless Network Server Log Data AWID Data Set Avaliable here: //a href="https://drive.google.com/open?id=0ByArNbFEPXXxZGo1bDJfSGRpbms">AWID: Wireless Network Server Log Data (1 GB zip) Paper that describe the data set: //a href="draft-Intrusion-Detection-in-802-11-Networks-Empirical-Evaluation-of-Threats.pdf">Paper: draft Intrusion Detection in 802-11 Networks Empirical Evaluation of Threats NSL-KDD data set for AnomalyDetection: //a href="https://www.kaggle.com/hassan06/nslkdd"> Kaggle: KDD data set for AnomalyDetection //a href="https://www.unb.ca/cic/datasets/nsl.html"> Data set Related Research site on AnomalyDetection For NIJ Challenging: Hot spot Analysis: See Clustering Section for more detailed information See More Info for NIJ Challegeing in the Clustering Section Download New data sets and Submission files of Winning Teams below to Learn from the NIJ (National Institute of Justice) Challenege site (Scroll down to see the links) Find Out How to Measure Hot Spots in score type PAI and PEI* for Every Category, Crime Type, time Frame. GIS Data Visualization API: Some of Research Projects Done in Big Data Analytics Lab: //a href="ResearchPoster_EECS_SunnieChung_MikeDArcy_UtkarshPatel.pdf"> Research Project: Document Clustering Using Word2Vec and Paragraph Vector //a href="ResearchPosters_NLPTeam.pdf">Research Project: Machine Learning Based Full Text Intelligent Search Engine //a href="Document Search EngineCollaborativeContentManagement.pdf">Research Project: Document Search Engine //a href="ResearchPoster_EECS_SunnieChung_DanielleAring.pdf">Research Project: Sentiment Analysis on Social Network Twitter //a href="JJLiu_Poster_FullDescription0914.pdf">Research Project: Feature Extraction and Selection Using AWID Data Project Presentations Examples Sentiment Analysis on Yelp Review Data //a href="CIS660ProjectSentimentAnalysisTwitterData.pdf"> Project: Text Mining for Sentiment Analysis: Predicting Review Stars (1 - 5) from Yelp Review Text data //a href="SummaryTextAnalysis_yelp_GunpingYu.pdf">Project: Yelp Challenge with Text Mining: Predicting Review Stars (1 - 5) from Yelp Review Text data //a href="TutorialHowtoGetFaceBookGraphAPIDataFacePager.pdf">How to Get FaceBook Graph API Data to Analytics with Hive Research Papers Based on: //a href="Sentiment_Pang_2002.pdf">Research Paper: Sentiment Analysis Using Classification //a href="SentimentAnalysis_Pres_Pang_2003Paper_XioadonLiu.pdf">Research Paper Presentation on Sentiment Analysis Using Classification (by Xiaodan Liu) //a href="SentimentPingConelle_2004.pdf">Research Paper: Sentiment Analysis Using Subjectivity Text Analysis Using Classification //a href="ISISTwitterVishnuSantosh.pdf">Text Mining on Twitter Data on ISIS terrorists group and the fan groups of ISIS Santosh Tankala, Vishnu Vishnuteja Thummanapelli, and Akhi Reddy Laxmanagari //a href="ISISTwitterVishnuSantosh.pdf">Tutorial: How To get Twitter Data Paper Based On: //a href="PaperGraphBasedHashTagSentinalAnalysis_SantoshV.pdf">Research Paper: Topic sentiment analysis in twitter: a graph-based hashtag sentiment classification approach Another Projects: Sentiment Analysis of Online Reviews: 2. Recommendation System on Amazon Product Data set //a href="CollaborativeFilteringSuhua.pdf"> Project: Collaborative Filtering Algorithm Implementation for for Recommendation System Suhua Wei //a href="RecommendationSystem_SagarDahiwala.pdf">Project: Recommendation System Using Amazon Customer Product Data Set for Collaborative Filtering Sagar Dahiwala Research Papers Based on: //a href="https://eecs.csuohio.edu/~sschung/CIS601/Amazon-Recommendations.pdf">Research Paper: Collaborative Filtering for Amazon Recommendation System /a> //a href="ProjectReport.v01Stanford_RecommendationSystem.pdf">Paper: Recommendation System from Stanford Project Report //a href="https://eecs.csuohio.edu/~sschung/CIS601/presentation01_DIzadnegahda.pdf">Paper PPresentation: Amazon Recommendation System for Collaborative Filtering 3. AnomalyDetection Using Classification: //a href="CIS660Project_IntrusionDetectionAnamalyDetectionNetworkServerData_Yanan.pdf">Intrusion Detection AnomalyDetection on Network Server Data by Sean Riehl and Yanan Lyu 4. Question Answering Systems: < href="QASystemImageTextSagarMyur.pdf"> Project: Question Answering System on Image and Text Data Mayur and Sagar 5. Clustering //a href="Arida_IndependentResearchSummary.pdf"> Independent Study Project: Outlier Detection on AWID Data Set Ahmad Arida //a href="Project_V2_AhmadDa.pdf">Project: Outlier Detection on NASA HTTP Logs Ahmad Arida //a href="NIJClusteringHotSpotAnalysisusingPython.pdf"> Project: Clustering and Hot Spot Analysis on Portland Map with NIJ Data Sarvesh Chande 6. Data Analytics on Social Media Data //a href="Data Mining In LinkedIn_final_ProjectOnlySlidesDanielleNoreen.pdf">Project: Data Mining over LinkedIn Data Danielle Aring and Noreen Halley //a href="SocialMediaProvenanceDataCollector_ResearchPaperSlidesDanielleNorren.pdf">Research Paper: A Tool for Collecting Provenance Data in Social Media Research Paper Presentation Guide (Extra Credit) For the extra credit Research paper presentation, you have to read and summarize the main method of the paper from the best conference sites in the Project List. Submit the slides in pptx and the paper. If you don't summarize the main content of the paper properly, no grade will be given. How to Read and What are the Important Contents to Summarize: //a href="http://cis.csuohio.edu/~sschung/CIS601/Research%20Paper%20101.pdf"> How to read and present a research paper //a href="https://eecs.csuohio.edu/~sschung/CIS601/Heideloff_Presentation2_PDWQO.pdf"> Example of Research Paper Presentation Slides //a href="https://eecs.csuohio.edu/~sschung/CIS601/CIS601IDS.html#Class_Notes"> More Research Paper Presenation examples |
| Class | Chapter / Topic / Specific Objectives / Activities |
| 1 |
Introduction to Big Data Analytics: |
| 1-4 |
|
| 4-5 |
|
| 6-10 |
|
| 10 |
|
| 11 |
|
| 11 |
|
| 11-12 |
|
| 14 |
|
| 15 |
|
| 13 |
|
| 16 |
|
==> Completion of Homeworks/Labs is required for obtaining a passing grade.
| This
is a tentative scale and |
Letter |
Quality Points |
|
||
| A |
> 93% |
A: Outstanding (student's performance is genuinely excellent) | |||
| A- |
90% - 93% |
||||
| B+ |
87% - 90% |
||||
| B |
82% - 87% |
B: Very Good (student's performance is clearly commendable but not necessarily outstanding) | |||
|
|
B- |
80% - 82% |
|||
|
|
C |
75% - 80% |
C: Good (student's performance meets every course requirement and is acceptable; not distinguished) | ||
| D | 65%-75% | D: Below Average (student's performance fails to meet course objectives and standards) | |||
|
|
F |
<65% |
F: Failure (student's performance is unacceptable) | ||
|
ADA Adherence. If you need course adaptations or accommodations because of a disability, if you have emergency medical information to share with me, or if you need special arrangements in case the building must be evacuated, please make an appointment with me as soon as possible. My office location and hours are listed on top of this syllabus. If you need further information, please contact the ACCESS office, phone number 687-5106. |
Programming standards
|
|
|