Data Analytics Codes
Klausur
Klausur
-
- 1 / 65
-
Flashcards
What is the Python code to generate a lift chart?
import kds as kds
kds.metrics.plot_lift(valid_y, predict_valid)
What is the Python code to make a confusion matrix?
predict_valid = logit_reg.predict(valid_X) cm2 = confusion_matrix(valid_y, predict_valid)
ConfusionMatrixDisplay(cm2).plot()
X = banking_df[["Income", "Family", "CCAvg", "Education", "Age"]] X = pd.get_dummies(X, prefix_sep="_", drop_first=True) Y = banking_df["has_mortgage"]# Data partitioningtrain_X, valid_X, train_y, valid_y = train_test_split(X, Y, test_size=0.4, random_state=10)# Logistic Regression logit_reg = LogisticRegression(solver="liblinear")logit_reg.fit(train_X, train_y)
What is the Python code to add explanatory variables and estimate it again?
X_full = banking_df[["Income", "Family", "CCAvg", "Education", "Age"]] X_full = pd.get_dummies(X_full, prefix_sep="_", drop_first=True)
X_full = X_full.astype(float) # Make sure that all columns have numerical data types
Y_full = banking_df["has_mortgage"] X_full = sm.add_constant
(X_full)logit_full_mod = sm.Logit(Y_full, X_full)
logit_full_mod_res = logit_full_mod.fit()print(logit_full_mod_res.summary())
What is the Python code to estimate a logit model: log(odds(has.mortgage = 1| income) = ß0 + ß1 * income?
X_simple = banking_df["Income"]
Y_simple = banking_df["has_mortgage"]
X_simple = sm.add_constant
(X_simple)logit_simple_mod = sm.Logit
(Y_simple, X_simple)logit_simple_mod_res = logit_simple_mod.fit()print(logit_simple_mod_res.summary())
What is the Python code to generate a new variable that takes the value 0 when Mortgage has the value 0 and takes the value 1 in all other cases?
banking_df["has_mortgage"] = [0 if x == 0 else 1 for x in banking_df["Mortgage"]]
banking_df.head()
What is the Python code to convert a variable into a categorical variable?
banking_df["Education"].value_counts().sort_index()
banking_df["Education"] = banking_df["Education"].map({1: "Undergrad", 2: "Graduate", 3: "Advanced/Professional"})
banking_df.head()
What is the Python code to replace the spaces in all variable names with underscores _?
banking_df.columns = [s.strip().replace(" ", "_") for s in banking_df.columns] banking_df.head()
What is the Python code to show the regression statistics of validation data?
print('Performance Measures (Validation data)') regressionSummary(valid_y, toyota_ml.predict(valid_X))
What is the Python code to show the regression statistics of training data?
print('Performance Measures (Training data)') regressionSummary(train_y, toyota_ml.predict(train_X))
What is the Python code for regression statistics?
# Fuel_Type transform in Dummies
X = toyota_df[['Fuel_Type', 'HP']]
y = toyota_df[['Price']]# Transform Fuel_Type in dummies
X = pd.get_dummies(X, drop_first=True)# Split the datatrain_X, valid_X, train_y,
valid_y = train_test_split(X, y, test_size=0.4)# Model
fittingtoyota_ml = LinearRegression()toyota_ml.fit(train_X, train_y)
What is the Python code for an OLS Regression to appreciate the influence of a variable based on another variable?
modg_X = toyota_df[['Fuel_Type']
]modg_X = pd.get_dummies(modg_X, drop_first=True)
modg_X = sm.add_constant(modg_X)
modg_X = modg_X.astype(float) # Make sure that all columns have numerical values# Model estimation and results
modg = sm.OLS(toyota_df['Price'], modg_X)res = modg.fit()print(res.summary())
What is the Python code to visualize the relationship between the selling price and the type of fuel in a stripplot?
with pd.option_context('mode.use_inf_as_na', True): sns.set(rc={'figure.figsize':(10,8), "figure.dpi":300,})
sns.set_theme(style="whitegrid")sns.stripplot(x="Fuel_Type", y="Price", data=toyota_df)
What is the Python code to visualize the relationship between the selling price and the type of fuel in a swarmplot?
with pd.option_context('mode.use_inf_as_na', True): sns.set(rc={'figure.figsize':(13,5), "figure.dpi":300,})
sns.set_theme(style="whitegrid")sns.swarmplot(x="Fuel_Type", y="Price", data=toyota_df, size=4)
What is the Python code to visualize the relationship between the selling price and the type of fuel in a boxplot?
sns.boxplot(x="Fuel_Type", y="Price", data=toyota_df, whis=100)
What is the Python code to calculate the arithmetic mean of a variable price for each category of another variable?
toyota_df.groupby('Fuel_Type').Price.mean()
What is the Python code for a frequency table of a variable (Fuel_Type)?
toyota_df.Fuel_Type.value_counts() banking_df.Mortgage.value_counts().sort_index()
What is the Python code to show two variables of a table?
toyota_df[["Fuel_Type", "Price"]]
What is the Python code for reading data with the encoding ISO-8859-1?
toyota_df = pd.read_csv('ToyotaCorolla.csv', encoding="ISO-8859-1") toyota_df.head()
Python code: Relationship between cut-off value and error rate and accuracy in a common plot?
ax = summary.plot(x="Cutoff", y="Accuracy", legend=False)
ax2 = ax.twinx() summary.plot(x="Cutoff", y="Error rate",
ax=ax2, legend=True, color="r")
What is the Python code to plot the evolution of a rate (lineplot)?
sns.lineplot(data=summary, x="Cutoff", y="Error rate")
What is the Python code for a table?
summary = pd.DataFrame({"Cutoff": [a, b, c], "Error rate": [a, b, c], "Accuracy": [a, b, c]}) summary
What is the Python code for a confusion matrix with a cutoff of 0.5?
predicted = ['owner' if p > 0.5 else 'nonowner' for p in owner_df.Probability]
classificationSummary(owner_df.Class, predicted, class_names=['nonowner', 'owner'])
errorrate50 = (1 + 2) / (10 + 2 + 1 + 11)
accuracy50 = 1 - errorrate50 sens50 = 11 / (11 + 1)
spec50 = 10 / (10 + 2)
print(f"Error rate: {errorrate50:4.3f}")
print(f"Accuracy: {accuracy50:4.3f}")
print(f"Sensitivity: {sens50:4.3f}")
print(f"Specificity: {spec50:4.3f}")
What is the code to cross-tabulate two binary variables?
pd.crosstab(housing_df["VARIABLE1"], housing_df["VARIABLE2"])
What is the code for a color-coded scatterplot according to VARIABLE 3 (binary) with sentences as outcomes?
sns.scatterplot(data=housing_df, y="VARIABLE1", x="VARIABLE2", hue="VARIABLE3_lb", alpha=0.8)
What is the code to name the binary outcome (0, 1) into sentences?
housing_df["VARIABLE_lb"] = housing_df["VARIABLE"].map({0: "Unter", 1: "Over"})
What is the code to show a frequency table of a binary variable?
housing_df.VARIABLE.value_counts()
What is the code for the standard deviation of a variable?
housing_df.VARIABLE.std()
What is the code for the maximum of a variable?
housing_df.VARIABLE.max()
What is the code for the minimum of a variable?
housing_df.VARIABLE.min()
What is the code for a color-coded scatterplot according to VARIABLE3 (binary)?
sns.scatterplot(data=housing_df, y="VARIABLE1", x="VARIABLE2", hue="VARIABLE3", alpha=0.8)
What is the code for a scatterplot with a regression line in it?
sns.regplot(data=housing_df, y="VARIABLE1", x="VARIABLE2")
What is the code for the scatterplot in Python using seaborn and making the dots a bit transparent?
sns.scatterplot(data=housing_df, y="VARIABLE1", x="VARIABLE2", alpha=0.8)
What is the Python code to plot two barcharts showing the mean?
sns.barplot(data=housing_df, y="VARIABLE1", x="VARIABLE2", estimator=mean)
What is the Python code to plot two boxplots without whiskers?
sns.boxplot(y=housing_df["VARIABLE1"], x=housing_df["VARIABLE2"], whis=[0, 100])
What is the Python code to plot two boxplots with whiskers?
sns.boxplot(y=housing_df["VARIABLE1"], x=housing_df["VARIABLE2"])
What is the Python code to plot a histogram?
sns.histplot(data=housing_df, x="VARIABLE", binwidth=10) # 'binwidth' can be adjusted or omitted depending on the task
What is the Python code to rename one variable?
housing_df = housing_df.rename(columns={"VARIABLE": "VARIABLENEW"})
trainData, temp = train_test_split(Housing_df, test_size=0.5, random_state=1)validData, testData = train_test_split(temp, test_size=0.4, random_state=1)print("Training: ", trainData.shape)print("Validation: ", validData.shape)print("Test: ", testData.shape)
What is the Python code to convert two categorical variables into dummy variables?
Housing_df = pd.get_dummies(Housing_df, columns=["Variable1", "Variable2"], prefix_sep="_", drop_first=True)
print(list(Housing_df.columns))Housing_df.head()