Tampilkan postingan dengan label dataset. Tampilkan semua postingan
Tampilkan postingan dengan label dataset. Tampilkan semua postingan

Senin, 20 April 2026

How to Construct a Panel Dataset from Scratch in R

 

Managing panel data in RStudio to estimate regression equations and determine the effect of independent variables on dependent variables.

One method for determining the influence of variables is using panel data. This type of influence allows us to estimate the dependent variable. Panel data is aggregated data in the form of

 

Creating a panel data structure

RStudio differs from other statistical software. To manage any analysis, it requires a data structure in RStudio. While other software simply copy-and-paste spreadsheet data, whether Excel or Google Sheets, to immediately manage the data, RStudio requires converting it into a data model recognized by RStudio.

The steps include importing a spreadsheet file and making some relatively simple adjustments to make your data easier to process. Specifically for panel analysis, the data structure required is pdata.frame, which is short for panel data frame. This differs from a regular data frame in RStudio because it considers both individual and time dimensions. This approach is what makes it different.

On the right, you can click "Import Data Set" and select Excel. There are several other options, such as SPSS, SAS, Stata, Text, and others. If you have a spreadsheet, select Excel.

After that, you will select multiple sheets. If you are working with multiple sheets in one file, you must select one of the sheets. Below that, you can select it. Therefore, you must pay attention to the neatness of your text. For example, if there is a gap between the table title and the data content, the empty table will be marked "NA" (Not Available), meaning the data is not available.

Preparing Excel as Data

To organize data, we can work with data. Because data with a spreadsheet is easier, we can organize it with data, as in the example below.

 




I uploaded the data in CSV format into RStudio.

tobinq3 <- read.csv2("~/jurnal/tobinq3.csv")

Then I can view the data like this:

View(tobinq3)

 

 


The data isn't in a pdataframe format yet, so we do it like this:

 

ptobinq=pdata.frame(tobinq3,index=c("Comp","Year"),drop.index = TRUE,row.names=TRUE)

 

The name ptobinq is the name I created to distinguish it from other files. From here, we've transformed the data structure into a panel dataframe. You'll see it look like this:

 

Classes ‘pdata.frame’ and 'data.frame':      40 obs. of  3 variables:

 $ DAR    : 'pseries' Named num  0.49 0.44 0.42 0.4 0.4 0.49 0.44 0.42 0.4 0.4 ...

  ..- attr(*, "names")= chr [1:40] "Adaro-2014" "Adaro-2015" "Adaro-2016" "Adaro-2017" ...

  ..- attr(*, "index")=Classes ‘pindex’ and 'data.frame':    40 obs. of  2 variables:

  .. ..$ Comp : Factor w/ 8 levels "Adaro","ATPK",..: 1 1 1 1 1 2 2 2 2 2 ...

  .. ..$ Tahun: Factor w/ 5 levels "2014","2015",..: 1 2 3 4 5 1 2 3 4 5 ...

 $ DER    : 'pseries' Named num  0.97 0.78 0.72 0.67 0.66 0.97 0.78 0.72 0.67 0.66 ...

  ..- attr(*, "names")= chr [1:40] "Adaro-2014" "Adaro-2015" "Adaro-2016" "Adaro-2017" ...

  ..- attr(*, "index")=Classes ‘pindex’ and 'data.frame':    40 obs. of  2 variables:

  .. ..$ Comp : Factor w/ 8 levels "Adaro","ATPK",..: 1 1 1 1 1 2 2 2 2 2 ...

  .. ..$ Tahun: Factor w/ 5 levels "2014","2015",..: 1 2 3 4 5 1 2 3 4 5 ...

 $ Tobin.Q: 'pseries' Named num  -0.2702 0.2346 0.2706 0.034 0.0336 ...

  ..- attr(*, "names")= chr [1:40] "Adaro-2014" "Adaro-2015" "Adaro-2016" "Adaro-2017" ...

  ..- attr(*, "index")=Classes ‘pindex’ and 'data.frame':    40 obs. of  2 variables:

  .. ..$ Comp : Factor w/ 8 levels "Adaro","ATPK",..: 1 1 1 1 1 2 2 2 2 2 ...

  .. ..$ Tahun: Factor w/ 5 levels "2014","2015",..: 1 2 3 4 5 1 2 3 4 5 ...

 - attr(*, "index")=Classes ‘pindex’ and 'data.frame':       40 obs. of  2 variables:

  ..$ Comp : Factor w/ 8 levels "Adaro","ATPK",..: 1 1 1 1 1 2 2 2 2 2 ...

  ..$ Tahun: Factor w/ 5 levels "2014","2015",..: 1 2 3 4 5 1 2 3 4 5 ..

 

It's clear that the word "dataframe" appears above. Then, there's the company index name and the year. This data is now ready to be converted into panel data analysis.

We can view the data at the top with the head command.

> head(ptobinq)
            DAR  DER     Tobin.Q
Adaro-2014 0.49 0.97 -0.27020301
Adaro-2015 0.44 0.78  0.23455470
Adaro-2016 0.42 0.72  0.27061008
Adaro-2017 0.40 0.67  0.03397098
Adaro-2018 0.40 0.66  0.03363631
ATPK-2014  0.49 0.97  0.30736531

 

The data appears to be different, so company and year are no longer variables as they are in Excel. With the dataframe model, both the time dimension and the company dimension are taken into account.



Kamis, 21 Agustus 2025

Mencari Dataset di Rstudio untuk latihan beberapa metode analisis statistik

 Ketika kita mau mencari data untuk pelatihan mungkin kita masih bingung. Atau kita ingin mencoba satu metode yang lain dalam statistik mungkin kita terbatas dengan data yang ada. Kebanyakan kita ini yang sebagai dosen seringnya hanaya metode yang umum saja seperti regresi. Kalau ada yang lebih canggih itu SEM. Atau metode yang paling mudah adalah uji non parametrik. 


Maka kalau anda mau untuk melakukan beberapa uji atau mau belajar membutuhkan data. data ini juga kebanyakan gratis dan bisa dipakai. Tetapi anda juga harus hati-hati mengenai lisensi data tersebut. Dengan adanya data ini anda akan bisa mendapatkan banyak data. 


Data yang sering dipakai seperti iris yakni data bunga iris. mtcars berisikan data mobil dan mesinnya. Data Lynx yang berisikan jumlah kucing lynx yang tertangkap di Kanada. dan banyak lainnya untuk mengeceknya anda bisa ketika perintah 


data() maka anda akan menemui data seperti Air Passengers, BJSales, BOD, CO2 DNase. Eurostock, Formalhyde dan banyak lagi. untuk menampiklan data anda bisa melakukan perintah seperti ini tulis langsung namanya datanya seperti langsung nama datanya yakni 

> data()
> iris
    Sepal.Length Sepal.Width Petal.Length
1            5.1         3.5          1.4
2            4.9         3.0          1.4
3            4.7         3.2          1.3
4            4.6         3.1          1.5
5            5.0         3.6          1.4
6            5.4         3.9          1.7
7            4.6         3.4          1.4
8            5.0         3.4          1.5
9            4.4         2.9          1.4
10           4.9         3.1          1.5

    Petal.Width    Species
1           0.2     setosa
2           0.2     setosa
3           0.2     setosa
4           0.2     setosa
5           0.2     setosa
6           0.4     setosa
7           0.3     setosa
8           0.2     setosa
9           0.2     setosa
10          0.1     setosa

sebanarnya datanya sampai 150. saya potong untuk menghemat pos kemudian kita bisa menulis perintah dibawah untuk melihat struktur data dari iris

> str(iris)
'data.frame':	150 obs. of  5 variables:
 $ Sepal.Length: num  5.1 4.9 4.7 4.6 5 5.4 4.6 5 4.4 4.9 ...
 $ Sepal.Width : num  3.5 3 3.2 3.1 3.6 3.9 3.4 3.4 2.9 3.1 ...
 $ Petal.Length: num  1.4 1.4 1.3 1.5 1.4 1.7 1.4 1.5 1.4 1.5 ...
 $ Petal.Width : num  0.2 0.2 0.2 0.2 0.2 0.4 0.3 0.2 0.2 0.1 ...
 $ Species     : Factor w/ 3 levels "setosa","versicolor",..: 1 1 1 1 1 1 1 1 1 1 ...
Kalau kita ingin melihat bagian awal saja dari tabel seperti menggunakan perintah head(data)
> head(iris)
  Sepal.Length Sepal.Width Petal.Length Petal.Width
1          5.1         3.5          1.4         0.2
2          4.9         3.0          1.4         0.2
3          4.7         3.2          1.3         0.2
4          4.6         3.1          1.5         0.2
5          5.0         3.6          1.4         0.2
6          5.4         3.9          1.7         0.4
  Species
1  setosa
2  setosa
3  setosa
4  setosa
5  setosa
6  setosa

KAlau kita ingin menampilkan bagian bawah kita bisa mengetikkan perintah tail(data)

> tail(iris)
    Sepal.Length Sepal.Width Petal.Length
145          6.7         3.3          5.7
146          6.7         3.0          5.2
147          6.3         2.5          5.0
148          6.5         3.0          5.2
149          6.2         3.4          5.4
150          5.9         3.0          5.1
    Petal.Width   Species
145         2.5 virginica
146         2.3 virginica
147         1.9 virginica
148         2.0 virginica
149         2.3 virginica
150         1.8 virginica
> View(iris)

iris maka akan muncul banyak nilai data. apalahi kalau data setnya kebanyakan. anda bisa memeriksa jenis data tersebut dengan perintah str(data) maka di situ akan menjelaskan data apa.misalnya dalam iris itu adalah data frame yang merupakan kumpulan dari data yang terdiri dari beberapa variabel yang di sana. Kemudian anda juga bisa menggunakan data dari package lain. kalau menggunakan package lain anda harus menjalankan perintah masuk ke package tersebut seperti dalam package MASS


>library(MASS)

>data(package="MASS)


maka anda akan mendapatkan banyak data di sini. ada beberapa package yang mempunyai data tertentu. sialahkan mencari data sesuai dengan kebutuhan anda, 

How to Construct a Panel Dataset from Scratch in R

  Managing panel data in RStudio to estimate regression equations and determine the effect of independent variables on dependent variables. ...