Tampilkan postingan dengan label regression. Tampilkan semua postingan
Tampilkan postingan dengan label regression. Tampilkan semua postingan

Senin, 20 April 2026

How to Construct a Panel Dataset from Scratch in R

 

Managing panel data in RStudio to estimate regression equations and determine the effect of independent variables on dependent variables.

One method for determining the influence of variables is using panel data. This type of influence allows us to estimate the dependent variable. Panel data is aggregated data in the form of

 

Creating a panel data structure

RStudio differs from other statistical software. To manage any analysis, it requires a data structure in RStudio. While other software simply copy-and-paste spreadsheet data, whether Excel or Google Sheets, to immediately manage the data, RStudio requires converting it into a data model recognized by RStudio.

The steps include importing a spreadsheet file and making some relatively simple adjustments to make your data easier to process. Specifically for panel analysis, the data structure required is pdata.frame, which is short for panel data frame. This differs from a regular data frame in RStudio because it considers both individual and time dimensions. This approach is what makes it different.

On the right, you can click "Import Data Set" and select Excel. There are several other options, such as SPSS, SAS, Stata, Text, and others. If you have a spreadsheet, select Excel.

After that, you will select multiple sheets. If you are working with multiple sheets in one file, you must select one of the sheets. Below that, you can select it. Therefore, you must pay attention to the neatness of your text. For example, if there is a gap between the table title and the data content, the empty table will be marked "NA" (Not Available), meaning the data is not available.

Preparing Excel as Data

To organize data, we can work with data. Because data with a spreadsheet is easier, we can organize it with data, as in the example below.

 




I uploaded the data in CSV format into RStudio.

tobinq3 <- read.csv2("~/jurnal/tobinq3.csv")

Then I can view the data like this:

View(tobinq3)

 

 


The data isn't in a pdataframe format yet, so we do it like this:

 

ptobinq=pdata.frame(tobinq3,index=c("Comp","Year"),drop.index = TRUE,row.names=TRUE)

 

The name ptobinq is the name I created to distinguish it from other files. From here, we've transformed the data structure into a panel dataframe. You'll see it look like this:

 

Classes ‘pdata.frame’ and 'data.frame':      40 obs. of  3 variables:

 $ DAR    : 'pseries' Named num  0.49 0.44 0.42 0.4 0.4 0.49 0.44 0.42 0.4 0.4 ...

  ..- attr(*, "names")= chr [1:40] "Adaro-2014" "Adaro-2015" "Adaro-2016" "Adaro-2017" ...

  ..- attr(*, "index")=Classes ‘pindex’ and 'data.frame':    40 obs. of  2 variables:

  .. ..$ Comp : Factor w/ 8 levels "Adaro","ATPK",..: 1 1 1 1 1 2 2 2 2 2 ...

  .. ..$ Tahun: Factor w/ 5 levels "2014","2015",..: 1 2 3 4 5 1 2 3 4 5 ...

 $ DER    : 'pseries' Named num  0.97 0.78 0.72 0.67 0.66 0.97 0.78 0.72 0.67 0.66 ...

  ..- attr(*, "names")= chr [1:40] "Adaro-2014" "Adaro-2015" "Adaro-2016" "Adaro-2017" ...

  ..- attr(*, "index")=Classes ‘pindex’ and 'data.frame':    40 obs. of  2 variables:

  .. ..$ Comp : Factor w/ 8 levels "Adaro","ATPK",..: 1 1 1 1 1 2 2 2 2 2 ...

  .. ..$ Tahun: Factor w/ 5 levels "2014","2015",..: 1 2 3 4 5 1 2 3 4 5 ...

 $ Tobin.Q: 'pseries' Named num  -0.2702 0.2346 0.2706 0.034 0.0336 ...

  ..- attr(*, "names")= chr [1:40] "Adaro-2014" "Adaro-2015" "Adaro-2016" "Adaro-2017" ...

  ..- attr(*, "index")=Classes ‘pindex’ and 'data.frame':    40 obs. of  2 variables:

  .. ..$ Comp : Factor w/ 8 levels "Adaro","ATPK",..: 1 1 1 1 1 2 2 2 2 2 ...

  .. ..$ Tahun: Factor w/ 5 levels "2014","2015",..: 1 2 3 4 5 1 2 3 4 5 ...

 - attr(*, "index")=Classes ‘pindex’ and 'data.frame':       40 obs. of  2 variables:

  ..$ Comp : Factor w/ 8 levels "Adaro","ATPK",..: 1 1 1 1 1 2 2 2 2 2 ...

  ..$ Tahun: Factor w/ 5 levels "2014","2015",..: 1 2 3 4 5 1 2 3 4 5 ..

 

It's clear that the word "dataframe" appears above. Then, there's the company index name and the year. This data is now ready to be converted into panel data analysis.

We can view the data at the top with the head command.

> head(ptobinq)
            DAR  DER     Tobin.Q
Adaro-2014 0.49 0.97 -0.27020301
Adaro-2015 0.44 0.78  0.23455470
Adaro-2016 0.42 0.72  0.27061008
Adaro-2017 0.40 0.67  0.03397098
Adaro-2018 0.40 0.66  0.03363631
ATPK-2014  0.49 0.97  0.30736531

 

The data appears to be different, so company and year are no longer variables as they are in Excel. With the dataframe model, both the time dimension and the company dimension are taken into account.



Jumat, 26 September 2025

Least Square Analysis For time series data

time series least square

time series least square

Least Square

One of time series method analysis is least square. The method similar to like least square that we use in other least square. in other word we use the time series data to analysis the trend of the time series data. The result of the method we can get the equation y = ax+b.

The analysis of the data

Prepare the time series data.

You can First, we collect the time series data and write in spreadsheet. Then, we upload the spreadsheet file to rstudio. we got a data frame data. we can run the test wit the data frame. but, if yo have already time series data you have to make adjustment the data.

# the example of making time series data
# Data from Rstudio (built-in) di R
data_lynx <- ts(lynx, start=c(1821), frequency=1)
# create time variable
timel <- as.numeric(time(data_lynx))

After set the data we examine least square method to the modified lynx data. the command is lm. you have to run it. After use the command, also use summary examine and you get the

# Examine the regressionm model 
tslynx <- lm(data_lynx ~ timel)
summary(tslynx)

Call:
lm(formula = data_lynx ~ timel)

Residuals:
   Min     1Q Median     3Q    Max 
 -1594  -1211   -755   1032   5366 

Coefficients:
             Estimate Std. Error t value Pr(>|t|)
(Intercept) -4630.034   8493.112  -0.545    0.587
timel           3.285      4.523   0.726    0.469

Residual standard error: 1589 on 112 degrees of freedom
Multiple R-squared:  0.004689,  Adjusted R-squared:  -0.004198 
F-statistic: 0.5276 on 1 and 112 DF,  p-value: 0.4691

the result is not satisfying. the F test show insignificant result of the model, meaning the least square cannot explain the relation between data and the time. We also see the independent variable (time) does not affect the dependent variable, (lynx). as we guess first, the least square may not satisfy to analyse long time series data. we also consider to use other method for forecasting.

##Other Test

After see the result we also need to pass some test ti make sure the model is good. Certainly, i do not have to continue the test, since the F Value is so bad. I can find other data to explain the least square methods.

as the common we also run the normality test for this method or least square. we can use kolmogorov test for residual for regression model. this formula is also consider the mean and also standard devation of data.

#run normality test
ks.test(residuals(tslynx), "pnorm", mean = mean(residuals(tslynx)), sd = sd(residuals(tslynx)))

    Asymptotic one-sample Kolmogorov-Smirnov test

data:  residuals(tslynx)
D = 0.20038, p-value = 0.0002115
alternative hypothesis: two-sided

The residuals of the test is not good. the p value is far below 0,05, meaning the distribution is not normal. we can not reject the null hypothesis that the data is not normally distributed. Though the result not good. we continue to other test such as autokorelasi and heteroskedaticity. we employ dw test for detecting autocorrelation and bp test for detecting heteroskedaticty. before we use the coman, we have to library zoo and library lmtest . we rund both test comand to the tslynx model.

library(zoo)

Attaching package: 'zoo'
The following objects are masked from 'package:base':

    as.Date, as.Date.numeric
library(lmtest)
dwtest(tslynx)

    Durbin-Watson test

data:  tslynx
DW = 0.56312, p-value = 2.308e-15
alternative hypothesis: true autocorrelation is greater than 0
bptest(tslynx)

    studentized Breusch-Pagan test

data:  tslynx
BP = 0.0009146, df = 1, p-value = 0.9759

The result of dwtest shows the autorocrelation due the p value of dwtest is below 0,05. tge result cannot reject null hypothesis that there is autocorellation in this model. While, the thest od Breuch Pagan test show that the model free from heteroskedaticty.

I think to cut the term of the long of data to make the term of data. perhaps i will consider to cut data. the lynx data is a yearly data from 1821 t0 1934. there is 114 data. perhaps i will use the data from 1900 to 1934 only.

picture Angela from Pixabay

How to Construct a Panel Dataset from Scratch in R

  Managing panel data in RStudio to estimate regression equations and determine the effect of independent variables on dependent variables. ...